Data processing method, data processing device, electronic device and storage medium

By sending the attribute data in the cached database to the distributed message queue when the timing task is triggered, and writing it to the search engine cluster after processing by consumers, the problem of inefficiency in writing of search engines under high data volume and frequent updates is solved, and quasi-real-time data writing and server resource saving is achieved.

CN114780564BActive Publication Date: 2025-05-16JINGDONG TECH HLDG CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210432671.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-05-16
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

When the amount of business data is large and updates are frequent, the data writing of search engines will occupy a large amount of computing resources, resulting in low writing efficiency.

Method used

By setting a timing task, when the timing task is triggered, the attribute data stored in multiple key values ​​in the cache database is sent to the distributed message queue. After receiving the message data from the message queue, the consumer processes the attribute data contained in the message data, and writes the processed data to the storage device of the search engine cluster.

Benefits of technology

It realizes secondary distribution and quasi-real-time writing of data, overcomes the underlying implementation limitations of search engines, improves data storage efficiency, and saves server resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114780564B_ABST
    Figure CN114780564B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method, including: in response to triggering a timed task, obtaining multiple key values ​​configured in a cache database; generating message data based on the key value and the attribute data stored in the key value for each key value; and sending the multiple message data to a distributed message queue, so that after receiving the message data from the distributed message queue, the consumer processes the attribute data contained in the message data and writes the processed attribute data into a storage device of a search engine cluster. In addition, the present disclosure also provides a data processing device, an electronic device, and a readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of big data technology, and more specifically, to a data processing method, a data processing device, an electronic device, a readable storage medium, and a computer program product. Background Art

[0002] Most of the business data generated by enterprises in the production and operation process are multi-dimensional data. For example, customer data usually includes basic information and relationship information of customers, such as customer accounts, sources, and types of belonging. In order to realize the search function of business data, in related technologies, enterprises usually store customer information in the nested model of the search engine.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art: due to the limitations of the underlying implementation of the search engine, when the amount of business data is large and updated frequently, the writing of data will occupy a large amount of computing resources. Summary of the invention

[0004] In view of this, the present disclosure provides a data processing method, a data processing device, an electronic device, a readable storage medium, and a computer program product.

[0005] One aspect of the present disclosure provides a data processing method, including: in response to triggering a scheduled task, obtaining multiple key values ​​configured in a cache database; generating message data for each key value based on the key value and the attribute data stored in the key value; and sending the multiple message data to a distributed message queue, so that after receiving the message data from the distributed message queue, the consumer processes the attribute data contained in the message data and writes the processed attribute data to the storage device of the search engine cluster.

[0006] According to an embodiment of the present disclosure, the above method also includes: obtaining change data generated in the database within a preset time period, wherein the above preset time period includes the trigger interval of the above scheduled task; for each change data obtained, determining a target key value from multiple key values ​​of the above cache database; and storing the above change data as attribute data of the above target key value in the above cache database.

[0007] According to an embodiment of the present disclosure, obtaining the change data generated in the database within the preset time period includes: obtaining log information generated in the database within the preset time period; and parsing the log information to obtain the change data.

[0008] According to an embodiment of the present disclosure, the above-mentioned change data is configured with a business line identifier and a customer number; and the multiple key values ​​of the above-mentioned cache database respectively belong to key value groups of multiple business lines.

[0009] According to an embodiment of the present disclosure, for each acquired change data, a target key value is determined from multiple key values ​​of the cache database, including: determining a target key value group based on a business line identifier of the change data; and determining the target key value from multiple key values ​​of the target key value group based on a customer number of the change data.

[0010] According to an embodiment of the present disclosure, after receiving the above-mentioned message data from the above-mentioned distributed message queue, the above-mentioned consumer processes the attribute data contained in the above-mentioned message data, and writes the processed attribute data into the storage device of the search engine cluster, including: after the above-mentioned consumer receives the above-mentioned message data, in response to successfully acquiring the distributed lock, taking out a preset number of attribute data from the attribute data contained in the above-mentioned message data to obtain multiple first target attribute data; processing the multiple above-mentioned first target attribute data to obtain second target attribute data; writing the above-mentioned second target attribute data into the storage device of the above-mentioned search engine cluster through a preset data interface; and releasing the above-mentioned distributed lock.

[0011] According to an embodiment of the present disclosure, the above-mentioned attribute data includes main data and sub-data of multiple dimensions, and the attribute data stored in the same key value has a preset order; the above-mentioned processing of the multiple first target attribute data includes: based on the above-mentioned preset order, deduplicating the attribute data having the same main data and sub-data in the multiple first target attribute data to obtain multiple third target attribute data; and for each third target attribute data, processing the sub-data of the above-mentioned third target attribute data based on the main data of the above-mentioned third target attribute data.

[0012] According to an embodiment of the present disclosure, for each third target attribute data, the sub-data of the third target attribute data is processed based on the master data of the third target attribute data, including: for each third target attribute data, when the master data of the third target attribute data is determined to be master data of a newly added operation type, the sub-data in the third target attribute data is deleted, and the sub-data of all dimensions corresponding to the master data is supplemented from the database; when the master data of the third target attribute data is determined to be master data of a modified operation type, the dimensions of the sub-data of the third target attribute data are retained, and the sub-data of all dimensions corresponding to the master data are supplemented from the database; when the third target attribute data does not contain master data, the dimensions of the sub-data of the third target attribute data are retained, the master data is supplemented from the database based on the business line identifier and customer number of the third target attribute data, and the sub-data of all dimensions corresponding to the master data are supplemented from the database.

[0013] According to an embodiment of the present disclosure, the above method also includes: after taking out a preset number of attribute data from the attribute data contained in the above message data, if the above message data still contains attribute data, sending the message data that completes the retrieval operation to the above distributed message queue, so that the above consumer can consume the message data that completes the retrieval operation again.

[0014] Another aspect of the present disclosure provides a data processing device, including: a first acquisition module, used to acquire multiple key values ​​configured in a cache database in response to triggering a scheduled task; a generation module, used to generate message data for each key value based on the key value and the attribute data stored in the key value; and a first processing module, used to send the multiple message data to a distributed message queue, so that after receiving the message data from the distributed message queue, the consumer processes the attribute data contained in the message data and writes the processed attribute data to the storage device of the search engine cluster.

[0015] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more instructions, wherein when the one or more instructions are executed by the one or more processors, the one or more processors implement the method as described above.

[0016] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above method when executed.

[0017] Another aspect of the present disclosure provides a computer program product, which includes computer executable instructions, and when the instructions are executed, are used to implement the method as described above.

[0018] According to the embodiments of the present disclosure, by setting a scheduled task, when the scheduled task is triggered, the attribute data stored in multiple key values ​​in the cache database can be sent to the distributed message queue, so that consumers can consume the message data in the message queue and write it to the storage device of the search engine. In the above technical means, by using a cache database, secondary distribution and quasi-real-time writing of data are realized, which at least partially overcomes the technical problem existing in the related art that due to the limitations of the underlying implementation of the search engine, when the amount of customer information is large and updated frequently, the writing of data will occupy a large amount of computing resources, thereby effectively improving data storage efficiency and saving server resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0020] Figure 1 An exemplary system architecture to which the data processing method according to an embodiment of the present disclosure can be applied is schematically shown.

[0021] Figure 2 The flowchart of the data processing method according to the embodiment of the present disclosure is schematically shown.

[0022] Figure 3 The diagram schematically shows a processing flow of the first distribution of change data according to an embodiment of the present disclosure.

[0023] Figure 4 The diagram schematically shows the processing flow of the second distribution of change data according to an embodiment of the present disclosure.

[0024] Figure 5 The block diagram schematically shows a data processing device according to an embodiment of the present disclosure.

[0025] Figure 6 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0027] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0028] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0029] In the case of using expressions such as "at least one of A, B, and C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0030] Elasticsearch is an open source search engine based on Apache Lucene(TM). With the high performance and full-featured features of Apache Lucene, it is widely used in enterprise business systems to search, analyze and explore business data. On the other hand, since the underlying layer of Elasticsearch is implemented in Java, due to language limitations, it will occupy a lot of computing resources when data is frequently written, making it impossible to write quickly.

[0031] In view of this, the embodiments of the present disclosure use a cache database to perform secondary distribution of data, thereby achieving quasi-real-time and frequent writing of data in scenarios where business data has many dimensions and a huge amount of data in each dimension.

[0032] Specifically, the embodiments of the present disclosure provide a data processing method, a data processing device, an electronic device, a readable storage medium and a computer program product. The method includes: in response to triggering a timed task, obtaining multiple key values ​​configured in a cache database; generating message data based on the key value and the attribute data stored in the key value for each key value; and sending the multiple message data to a distributed message queue, so that after receiving the message data from the distributed message queue, the consumer processes the attribute data contained in the message data and writes the processed attribute data to the storage device of the search engine cluster.

[0033] In the technical solution disclosed in the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and do not violate public order and good morals.

[0034] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0035] Figure 1 The exemplary system architecture to which the data processing method according to the embodiment of the present disclosure can be applied is schematically shown. It should be noted that: Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0036] like Figure 1 As shown, the system architecture 100 according to this embodiment may include front-end devices 101 and 102 , a back-end server 103 and a server cluster 104 .

[0037] The front-end devices 101 and 102 may be various electronic devices that support human-computer interaction functions, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0038] Various client applications may be configured on the front-end devices 101 and 102, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients and / or social platform software.

[0039] The backend server 103 may be a server providing various services, and may be configured with a storage device, wherein the storage device includes a database and a cache database. Any operation performed by the user in the frontend devices 101 and 102 may be characterized as a data change in the database of the backend server 103 .

[0040] The server cluster 104 may be composed of multiple servers that provide the same service. For example, the server cluster 104 may be composed of multiple servers that provide support for a search engine.

[0041] The front-end devices 101, 102 and the back-end server 103, and the back-end server 103 and the server cluster 104 may be connected via a network, and the network may include various connection types, such as wired and / or wireless communication links.

[0042] It should be noted that the data processing method provided in the embodiment of the present disclosure can generally be executed by the server 103. Accordingly, the data processing device provided in the embodiment of the present disclosure can generally be set in the server 103. The data processing method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 103 and can establish communication with the front-end devices 101, 102, the server 103 and the server cluster 104. Accordingly, the data processing device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 103 and can establish communication with the front-end devices 101, 102, the server 103 and the server cluster 104.

[0043] For example, the user's operation in the front-end device 101 can be reflected through the network as a data change in the database of the server 103. The data change can be collected by the cache database in the server 103, and the data processing method provided by the embodiment of the present disclosure can be executed to process the changed data; or, other servers or server clusters can also read the changed data from the cache database of the server 103 and execute the data processing method provided by the embodiment of the present disclosure.

[0044] It should be understood that Figure 1 The number of front-end devices, servers and server clusters in the embodiment is only for illustration. According to the implementation requirements, any number of front-end devices, servers and server clusters may be provided.

[0045] Figure 2 The flowchart of the data processing method according to the embodiment of the present disclosure is schematically shown.

[0046] like Figure 2 As shown, the method includes operations S201 to S203.

[0047] In operation S201, in response to triggering a scheduled task, a plurality of key values ​​configured in a cache database are obtained.

[0048] In operation S202 , for each key value, message data is generated based on the key value and the attribute data stored in the key value.

[0049] In operation S203, multiple message data are sent to the distributed message queue, so that after receiving the message data from the distributed message queue, the consumer processes the attribute data contained in the message data and writes the processed attribute data into the storage device of the search engine cluster.

[0050] According to an embodiment of the present disclosure, a scheduled task may be a task executed at a certain frequency. For example, the execution frequency of a scheduled task may be set to 1 second, that is, the time window between two adjacent scheduled tasks is 1 second.

[0051] According to an embodiment of the present disclosure, the cache database may be any type of key-value database, such as redis, memcache, squid, etc.

[0052] According to an embodiment of the present disclosure, multiple key values ​​may be pre-configured in the cache database, and business data written into the cache database may be distributed to multiple key values ​​to avoid hot key problems caused by a single key value being accessed by a large number of requests in a short period of time.

[0053] According to an embodiment of the present disclosure, when a scheduled task is triggered, the data contained in each key-value pair in the cache database may be generated into a message data.

[0054] According to an embodiment of the present disclosure, the generated multiple message data may be sent to the distributed message queue in any order.

[0055] According to an embodiment of the present disclosure, the consumer may be any server in a search engine cluster.

[0056] According to an embodiment of the present disclosure, after consuming message data, consumers can obtain the data contained in the message data, that is, the corresponding key values ​​and attribute data in the original cache database, and process the obtained attribute data, including but not limited to deletion, merging, supplementation, etc.

[0057] According to the embodiments of the present disclosure, by using a cache database and a distributed message queue in conjunction, the frequency of writing data to the search engine's storage device in a scheduled task is effectively reduced, thereby also reducing the occupation of computing resources by the write operation; on the other hand, by setting a smaller scheduled task trigger interval, quasi-real-time writing of data can also be achieved.

[0058] According to the embodiments of the present disclosure, by setting a scheduled task, when the scheduled task is triggered, the attribute data stored in multiple key values ​​in the cache database can be sent to the distributed message queue, so that consumers can consume the message data in the message queue and write it to the storage device of the search engine. In the above technical means, by using a cache database, secondary distribution and quasi-real-time writing of data are realized, which at least partially overcomes the technical problem existing in the related art that due to the limitations of the underlying implementation of the search engine, when the amount of customer information is large and updated frequently, the writing of data will occupy a large amount of computing resources, thereby effectively improving data storage efficiency and saving server resources.

[0059] Reference below Figure 3-4 , combined with specific embodiments Figure 2 The method shown is further explained.

[0060] Figure 3 The diagram schematically shows a processing flow of the first distribution of change data according to an embodiment of the present disclosure.

[0061] like Figure 3 As shown, the processing flow includes operations S301 to S304.

[0062] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order of multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.

[0063] In operation S301, the change data generated in the database is obtained.

[0064] In operation S302, the changed data is distributed to the corresponding key value of the cache database.

[0065] In operation S303, determine whether the scheduled task is triggered; if the scheduled task is not triggered, return to execute operation S301 to continue storing the data generated in the database into the cache database; in response to triggering the scheduled task, execute operation S304.

[0066] In operation S304, each key value in the cache database is packaged into message data and sent to the message queue.

[0067] According to the embodiments of the present disclosure, the database may be a relational database such as MySQL, Oracle, etc., or a non-relational database such as MongoDB, CouchDB, etc., which is not limited here.

[0068] According to an embodiment of the present disclosure, the database implements the storage of business data generated in the business system. When business data is added, modified or deleted, the log file of the database can record the relevant information of the data change. For example, when any field in mysql changes, binlog will generate a log record based on the change of the field.

[0069] According to an embodiment of the present disclosure, the change data generated in the database can be obtained by obtaining the log information generated in the database and parsing the log information.

[0070] According to the embodiments of the present disclosure, the acquired change data can be sent to the cache database through the data change transmission system. With the help of the data change transmission system distributing data according to the id hash, it can be ensured that there is a sequence between multiple changes of a piece of data. For example, id=1, name=a, the name of the data is changed to b and c successively, and the data received by the cache database is also in the order of name=b and name=c.

[0071] According to an embodiment of the present disclosure, a target key value may be determined from a plurality of key values ​​pre-configured in a cache database, and then the changed data may be stored in the cache database as attribute data of the target key value.

[0072] According to an embodiment of the present disclosure, when configuring key values ​​for a cache database, a key value group may be configured for each business line, and each key value group may include multiple key values. For example, business line 100 is configured with 2 key values, which may be named redis_100_0 and redis_100_1 respectively.

[0073] According to an embodiment of the present disclosure, by configuring multiple key values ​​for each business line, it is possible to effectively prevent a hot key problem caused by excessive data volume in a certain business line, thereby affecting the reading and writing of the cache database.

[0074] According to an embodiment of the present disclosure, the change data may be configured with a business line identifier and a customer number. When the change data is allocated, the target key value group may be determined based on the business line identifier of the change data, and the target key value may be determined from multiple key values ​​of the target key value group based on the customer number of the change data. For example, if the business line identifier of the change data is 100 and the customer number is 100001, the change data may be allocated to the key value redis_100_1 according to the set rules.

[0075] According to an embodiment of the present disclosure, a scheduled task can be configured with a certain trigger interval. Within the time window of the trigger interval, the cache database continues to obtain change data from the database. After the scheduled task is triggered, the cache database can package the data obtained within the time window into message data and send the message data to a distributed message queue for consumption by consumers.

[0076] According to an embodiment of the present disclosure, all key values ​​configured in the cache database may be packaged as message data.

[0077] In some embodiments, the key values ​​for the packaging operation can also be filtered. If the corresponding key values ​​do not contain attribute data, a "no changed data" identifier can be added to the packaged message data so that consumers can directly discard the message data after receiving it, thereby improving the efficiency of data synchronization.

[0078] Figure 4 The diagram schematically shows the processing flow of the second distribution of change data according to an embodiment of the present disclosure.

[0079] like Figure 4 As shown, the processing flow includes operations S401 to S409.

[0080] In operation S401, it is determined whether the distributed lock is acquired successfully; if it is determined that the distributed lock is not acquired successfully, operation S402 is performed; if it is determined that the distributed lock is acquired successfully, operation S403 is performed.

[0081] In operation S402, the locking operation is exited.

[0082] In operation S403, message data is consumed from the distributed message queue to obtain a plurality of attribute data.

[0083] In operation S404, a preset number of attribute data are taken out from the plurality of attribute data to obtain first target attribute data.

[0084] In operation S405, the first target attribute data is deduplicated and supplemented to obtain second target attribute data.

[0085] In operation S406, the second target attribute data is written into a storage device of the search engine cluster through a preset data interface.

[0086] In operation S407, it is determined whether there is any remaining attribute data in the message data; if it is determined that there is still any remaining attribute data in the message data, operation S408 is performed; if it is determined that the message data is empty, operation S409 is performed.

[0087] In operation S408, the remaining attribute data is packaged into message data and sent to the distributed message queue.

[0088] In operation S409, the distributed lock is released.

[0089] According to the embodiments of the present disclosure, the implementation method of the distributed lock is not limited. For example, the distributed lock can be implemented based on redis, zookeeper, etc.

[0090] According to an embodiment of the present disclosure, by using a distributed lock, the atomicity of data writing operations in resources can be ensured to avoid logical errors in resources.

[0091] According to an embodiment of the present disclosure, after exiting the locking operation, according to the program logic configured by the developer, you can wait for a period of time before trying to lock again, or you can directly feedback the operation failure to the user and wait for the user's next operation instruction.

[0092] According to an embodiment of the present disclosure, the preset number can be set according to a specific business scenario, for example, it can be set to 1000.

[0093] According to an embodiment of the present disclosure, when the attribute data included in the key value is less than the preset quantity, the entire amount of attribute data in the key value may be retrieved.

[0094] According to an embodiment of the present disclosure, when the attribute data contained in the key value is greater than the preset quantity, the remaining attribute data in the key value can be repackaged into message data and sent to the message queue to be consumed by the next consumer.

[0095] According to an embodiment of the present disclosure, attribute data may include data of multiple dimensions, and data of one dimension may be set as primary data, and data of other dimensions may be set as sub-data. For example, in an application scenario of online shopping, customer identification or customer information in business data may be set as primary data, and order information may be set as sub-data.

[0096] According to an embodiment of the present disclosure, the specific steps of processing the first target attribute data may include:

[0097] First, single deduplication may be performed, and based on a preset order, attribute data with the same main data and sub-data in a plurality of first target attribute data may be deduplicated to obtain a plurality of third target attribute data.

[0098] According to an embodiment of the present disclosure, the preset order may be the original storage order of the attribute data.

[0099] According to an embodiment of the present disclosure, when a piece of data undergoes multiple changes, deduplication is performed based on a preset order, and only the data obtained after the last change can be retained to reduce the amount of data that needs to be subsequently processed and improve the efficiency of data synchronization.

[0100] Then, deduplication may be performed between the main data and the sub-data in each first target attribute data, and for each third target attribute data, the sub-data of the third target attribute data may be processed based on the main data of the third target attribute data.

[0101] According to an embodiment of the present disclosure, for each third target attribute data, when the master data of the third target attribute data is determined to be the master data of the newly added operation type, the sub-data in the third target attribute data is deleted, and the sub-data of all dimensions corresponding to the master data is supplemented from the database. For example, if the customer information represented by the master data is not recorded in the storage device of the current search engine, it can be considered that the customer corresponding to the customer information is a new user. In order to avoid missing the information of the new user, the sub-data corresponding to the master data can be deleted, and the full amount of data related to the user is obtained from the database as the sub-data of the master data.

[0102] In the case where the master data of the third target attribute data is determined to be the master data of the modification operation type, the dimensions of the sub-data of the third target attribute data are retained, and the sub-data of all dimensions corresponding to the master data are supplemented from the database. For example, if the sub-data corresponding to the master data is added or modified several times, only the data after the last change can be retained.

[0103] When the third target attribute data does not contain master data, the dimensions of the sub-data of the third target attribute data are retained, the master data is supplemented from the database based on the business line identifier and customer number of the third target attribute data, and the sub-data of all dimensions corresponding to the master data are supplemented from the database.

[0104] According to an embodiment of the present disclosure, the preset data interface may be set according to the search engine. For example, when the search engine is elasticsearch, the preset data interface may be a bulk api.

[0105] According to the embodiments of the present disclosure, by merging attribute data, invalid data in the original data can be effectively deleted, the number of write operations to the storage device of the search engine can be reduced, the computing pressure of the search engine server can be reduced, and server resources can be saved.

[0106] Figure 5 The block diagram schematically shows a data processing device according to an embodiment of the present disclosure.

[0107] like Figure 5As shown, the data processing device 500 includes a first acquisition module 510 , a generation module 520 and a first processing module 530 .

[0108] The first acquisition module 510 is used to acquire multiple key values ​​configured in the cache database in response to triggering a scheduled task.

[0109] The generating module 520 is used to generate message data for each key value based on the key value and the attribute data stored in the key value.

[0110] The first processing module 530 is used to send multiple message data to the distributed message queue so that after receiving the message data from the distributed message queue, the consumer processes the attribute data contained in the message data and writes the processed attribute data into the storage device of the search engine cluster.

[0111] According to the embodiments of the present disclosure, by setting a scheduled task, when the scheduled task is triggered, the attribute data stored in multiple key values ​​in the cache database can be sent to the distributed message queue, so that consumers can consume the message data in the message queue and write it to the storage device of the search engine. In the above technical means, by using a cache database, secondary distribution and quasi-real-time writing of data are realized, which at least partially overcomes the technical problem existing in the related art that due to the limitations of the underlying implementation of the search engine, when the amount of customer information is large and updated frequently, the writing of data will occupy a large amount of computing resources, thereby effectively improving data storage efficiency and saving server resources.

[0112] According to an embodiment of the present disclosure, the data processing device 500 further includes a second acquisition module, a first determination module and a first storage module.

[0113] The second acquisition module is used to acquire the change data generated in the database within a preset time period, wherein the preset time period includes the triggering interval of the scheduled task.

[0114] The determination module is used to determine a target key value from multiple key values ​​in the cache database for each acquired change data.

[0115] The storage module is used to store the changed data as attribute data of the target key value in the cache database.

[0116] According to an embodiment of the present disclosure, the second acquisition module includes a first acquisition unit and a second acquisition unit.

[0117] The first acquisition unit is used to acquire log information generated in a database within a preset time period.

[0118] The second acquisition unit is used to parse the log information to obtain the change data.

[0119] According to an embodiment of the present disclosure, the change data is configured with a business line identifier and a customer number, and multiple key values ​​of the cache database belong to key value groups of multiple business lines respectively.

[0120] According to an embodiment of the present disclosure, the determination module includes a first determination unit and a second determination unit.

[0121] The first determining unit is used to determine a target key value group based on a business line identifier of the changed data.

[0122] The second determining unit is configured to determine a target key value from a plurality of key values ​​of the target key value group based on the customer number of the changed data.

[0123] According to an embodiment of the present disclosure, the first processing module 530 includes a first processing submodule, a second processing submodule, a third processing submodule, and a fourth processing submodule.

[0124] The first processing submodule is used to extract a preset number of attribute data from the attribute data included in the message data in response to successfully acquiring the distributed lock after the consumer receives the message data, so as to obtain a plurality of first target attribute data.

[0125] The second processing submodule is used to process the plurality of first target attribute data to obtain second target attribute data.

[0126] The third processing submodule is used to write the second target attribute data into the storage device of the search engine cluster through a preset data interface.

[0127] The fourth processing submodule is used to release the distributed lock.

[0128] According to an embodiment of the present disclosure, the attribute data includes main data and sub-data of multiple dimensions, and the attribute data stored in the same key value has a preset order.

[0129] According to an embodiment of the present disclosure, the second processing submodule includes a first processing unit and a second processing unit.

[0130] The first processing unit is used to deduplicate attribute data having the same main data and sub-data in the plurality of first target attribute data based on a preset order to obtain a plurality of third target attribute data.

[0131] The second processing unit is used to process the sub-data of each third target attribute data based on the main data of the third target attribute data.

[0132] According to an embodiment of the present disclosure, the second processing unit includes a first processing sub-unit, a second processing sub-unit and a third processing sub-unit.

[0133] The first processing sub-unit is used to delete the sub-data in the third target attribute data for each third target attribute data, when the main data of the third target attribute data is determined to be the main data of the newly added operation type, and supplement the sub-data of all dimensions corresponding to the main data from the database.

[0134] The second processing sub-unit is used to retain the dimensions of the sub-data of the third target attribute data and supplement the sub-data of all dimensions corresponding to the main data from the database when the main data of the third target attribute data is determined to be the main data of the modification operation type.

[0135] The third processing sub-unit is used to retain the dimensions of the sub-data of the third target attribute data when the third target attribute data does not contain the master data, supplement the master data from the database based on the business line identifier and customer number of the third target attribute data, and supplement the sub-data of all dimensions corresponding to the master data from the database.

[0136] According to an embodiment of the present disclosure, the data processing device 500 further includes a second processing module.

[0137] The second processing module is used to, after taking out a preset number of attribute data from the attribute data contained in the message data, send the message data that has completed the taking out operation to the distributed message queue if the message data still contains attribute data, so that the consumer can consume the message data that has completed the taking out operation again.

[0138] According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits, or at least part of the functions of any one of them can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be at least partially implemented as hardware circuits, such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems on chips, systems on substrates, systems on packages, application specific integrated circuits (ASICs), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, submodules, units, and subunits can be at least partially implemented as computer program modules, and when the computer program modules are run, the corresponding functions can be performed.

[0139] For example, any multiple of the first acquisition module 510, the generation module 520, and the first processing module 530 can be combined in one module / unit / sub-unit for implementation, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the first acquisition module 510, the generation module 520, and the first processing module 530 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware or in a suitable combination of any of them. Alternatively, at least one of the first acquisition module 510 , the generation module 520 , and the first processing module 530 may be at least partially implemented as a computer program module, and when the computer program module is executed, a corresponding function may be performed.

[0140] It should be noted that the data processing device part in the embodiments of the present disclosure corresponds to the data processing method part in the embodiments of the present disclosure. The description of the data processing device part specifically refers to the data processing method part, which will not be repeated here.

[0141] Figure 6 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure is schematically shown. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0142] like Figure 6 As shown, the computer electronic device 600 according to the embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 to the random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include an onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.

[0143] In RAM 603, various programs and data required for the operation of electronic device 600 are stored. Processor 601, ROM 602 and RAM 603 are connected to each other via bus 604. Processor 601 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 602 and / or RAM 603. It should be noted that the program can also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.

[0144] According to an embodiment of the present disclosure, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage portion 608 as needed.

[0145] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.

[0146] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.

[0147] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include, but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.

[0148] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603 .

[0149] The embodiments of the present disclosure also include a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the data processing method provided by the embodiments of the present disclosure.

[0150] When the computer program is executed by the processor 601, the above functions defined in the system / device of the embodiment of the present disclosure are executed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0151] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0152] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).

[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. It can be understood by those skilled in the art that the features recorded in the various embodiments and / or claims of the present disclosure can be combined and / or combined in a variety of ways, even if such a combination or combination is not explicitly recorded in the present disclosure. In particular, without departing from the spirit and teaching of the present disclosure, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.

[0154] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above separately, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. The scope of the present disclosure is defined by the attached claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A data processing method, comprising: In response to triggering a scheduled task, multiple key values ​​configured in the cache database are obtained; For each key value, respectively, generating message data based on the key value and the attribute data stored in the key value; as well as The multiple message data are sent to a distributed message queue so that after receiving the message data from the distributed message queue, the consumer processes the multiple attribute data corresponding to the same key value contained in the message data, and writes the processed attribute data into the storage device of the search engine cluster, wherein the processing method includes deletion, merging or supplementation.

2. The method according to claim 1, further comprising: Acquire the change data generated in the database within a preset time period, wherein the preset time period includes the triggering interval of the scheduled task; For each acquired change data, determining a target key value from a plurality of key values ​​in the cache database; and The changed data is used as attribute data of the target key value and stored in the cache database.

3. The method according to claim 2, wherein: The step of obtaining the change data generated in the database within the preset time period includes: Acquiring log information generated in the database within the preset time period; and The log information is parsed to obtain the change data.

4. The method according to claim 2, wherein: The change data is configured with a business line identifier and a customer number; The multiple key values ​​of the cache database belong to key value groups of multiple business lines respectively.

5. The method according to claim 4, wherein: For each acquired change data, determining a target key value from a plurality of key values ​​of the cache database includes: Determining a target key value group based on the business line identifier of the changed data; and The target key value is determined from a plurality of key values ​​of the target key value group based on the customer number of the change data.

6. The method according to claim 4, wherein: After receiving the message data from the distributed message queue, the consumer processes the attribute data contained in the message data and writes the processed attribute data into a storage device of the search engine cluster, including: After the consumer receives the message data, in response to successfully acquiring the distributed lock, a preset number of attribute data are taken out from the attribute data included in the message data to obtain a plurality of first target attribute data; Processing a plurality of the first target attribute data to obtain second target attribute data; Writing the second target attribute data into the storage device of the search engine cluster through a preset data interface; and Release the distributed lock.

7. The method according to claim 6, wherein: The attribute data includes main data and sub-data of multiple dimensions, and the attribute data stored in the same key value has a preset order; The processing of the plurality of first target attribute data comprises: Based on the preset order, attribute data having the same main data and sub-data in the plurality of first target attribute data are deduplicated to obtain a plurality of third target attribute data; as well as For each third target attribute data, the sub data of the third target attribute data is processed based on the main data of the third target attribute data.

8. The method according to claim 7, wherein: For each third target attribute data, processing the sub-data of the third target attribute data based on the main data of the third target attribute data includes: For each third target attribute data, when the master data of the third target attribute data is determined to be the master data of the newly added operation type, the sub-data in the third target attribute data is deleted, and the sub-data of all dimensions corresponding to the master data is supplemented from the database; In the case where the master data of the third target attribute data is determined to be the master data of the modification operation type, retaining the dimensions of the sub-data of the third target attribute data, and supplementing the sub-data of all dimensions corresponding to the master data from the database; In the case that the third target attribute data does not contain master data, the dimensions of the sub-data of the third target attribute data are retained, the master data is supplemented from the database based on the business line identifier and customer number of the third target attribute data, and the sub-data of all dimensions corresponding to the master data are supplemented from the database.

9. The method according to claim 6, further comprising: After taking out a preset amount of attribute data from the attribute data contained in the message data, if the message data still contains attribute data, the message data on which the taking out operation is completed is sent to the distributed message queue so that the consumer can consume the message data on which the taking out operation is completed again.

10. A data processing device, comprising: A first acquisition module, configured to acquire multiple key values ​​configured in a cache database in response to triggering a scheduled task; A generating module, for generating message data for each key value based on the key value and the attribute data stored in the key value; as well as The first processing module is used to send the multiple message data to the distributed message queue, so that after receiving the message data from the distributed message queue, the consumer processes the multiple attribute data corresponding to the same key value contained in the message data, and writes the processed attribute data into the storage device of the search engine cluster, wherein the processing method includes deletion, merging or supplementation.

11. An electronic device, comprising: one or more processors; a memory for storing one or more instructions, Wherein, when the one or more instructions are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having executable instructions stored thereon, wherein when the executable instructions are executed by a processor, the processor is enabled to implement the method according to any one of claims 1 to 9.

13. A computer program product, comprising computer executable instructions, which are used to implement the method of any one of claims 1 to 9 when executed.

Citation Information

Patent Citations

  • Metadata retrieval method and device, storage medium and electronic equipment

    CN111858496A