Data processing method and device
By independent of offline data acquisition and loading, and adopting distributed processing and remote file center storage, the problem of coupling offline data and online services is solved, and the complete decoupling of offline and online is achieved, improving the stability of online services and the fault tolerance of data management is improved.
Patent Information
- Application Number
- CN202311551529.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
In the prior art, offline data is coupled with online services, resulting in large occupancy of online service resources, serious CPU jitter, affecting the stability of online services; the data source pressure is high and there is no fault tolerance, making data management difficult to control uniformly.
By independent of offline data acquisition and loading, it can independently undertake the function of offline data loading and acquisition, and configure multiple machines in distributed locking and horizontal expansion to perform data splitting and sharding processing. At the same time, the remote file center is used to store intermediate data and send notification messages through the message processing center, allowing the client to listen and callback data.
It realizes complete decoupling of offline and online, improves online service stability, reduces resource consumption and resource competition, reduces data source pressure, and improves the fault tolerance and version control capabilities of data management.
Smart Images

Figure CN120020755A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and particularly to a method and apparatus for data processing. Background Art
[0002] The production of offline data is generally written in the online service in the form of a scheduled task. The scheduled task is used to query and assemble data for data distribution and processing. Currently, in many cases on the operation side, data needs to be downloaded. However, this kind of download generally directly queries the data, performs data query and processing, and then provides a download link for downloading. In cases where the data volume is very large, or the query period is very long, or the SQL (Relational Database Management System) is slow, etc., it will not only cause the file export to fail, but may also cause the database CPU (Central Processing Unit) to soar, affecting the online business.
[0003] In the process of implementing the present invention, the inventor found that there are at least the following problems in the prior art:
[0004] Offline and online coupling: The offline service is coupled in the online service, which will occupy a large amount of resources of the online service, resulting in large fluctuations in the online service CPU / GC (Central Processing Unit / Garbage Collection), affecting the startup duration of the online service; high data source pressure: When obtaining offline data, it is necessary to schedule all online machines to obtain data from the data source simultaneously, which puts too much pressure on the data source; and data management fault tolerance problem: There is no version control ability for data, no fault tolerance ability, and the data is scattered and cannot be uniformly managed. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a method for data processing, which can separate offline data from the online service, relieve the pressure on the data source, streamline and configure the data production process, and improve the fault tolerance ability. The offline data acquisition and loading are independent into an independent service, and this service alone undertakes the function of offline data loading and acquisition.
[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, there is provided a method for data processing, including:
[0007] Pulling raw data and performing processing; and
[0008] Receiving the intermediate data obtained through the processing and writing the intermediate data into a remote file center, and sending a notification message with predetermined content related to the intermediate data to a message processing center, wherein the change of the predetermined content of the notification message can be monitored by a client.
[0009] Optionally, in the data processing method according to an aspect of the embodiment, the processing of the original data includes: in response to a user's demand, pulling the original data from the server and storing it in the memory, and processing the stored original data in the memory to generate the intermediate data.
[0010] Optionally, in the data processing method according to an aspect of the embodiment, receiving the intermediate data obtained through the processing and writing the intermediate data to the remote file center includes: serializing the full amount result of the intermediate data into a predetermined format, and generating a first temporary file locally; and storing the generated first temporary file in the remote file center and deleting the first temporary file locally.
[0011] Optionally, in the data processing method according to an aspect of the embodiment, the method is implemented by configuring multiple machines in a distributed lock and horizontally scalable manner, the multiple machines are redundantly deployed and use lock contention, and the original data is split into multiple data blocks, and each machine processes one data block to generate a corresponding sub-file and stores it in the remote file center.
[0012] Optionally, in the data processing method according to an aspect of the embodiment, it further includes: setting and monitoring the timeout for the pulling duration of the original data, the processing duration of the original data, and the writing duration, and setting timeout alarms.
[0013] To achieve the above object, according to an aspect of an embodiment of the present invention, there is also provided a data processing method, including:
[0014] Listening to a notification message with a predetermined content sent from the server to the message processing center, and listening to changes in the predetermined content of the notification message, where the predetermined content of the notification message is related to the intermediate data obtained through processing the original data; and
[0015] In response to the change in the predetermined content of the notification message, receiving callback information from the message processing center, and parsing the message information of the callback information to obtain data.
[0016] Optionally, in the data processing method according to one aspect of the embodiment, parsing the message information of the callback information to obtain data includes: obtaining the file download address of the file stored in the remote file center, downloading the file as a second temporary file to the local, deserializing the second temporary file into a target object in memory, deleting the second temporary file locally, and verifying the data in the target object and the callback message, and when the verification passes, triggering to replace the old target object in memory with the target object.
[0017] Optionally, in the data processing method according to one aspect of the embodiment, the file is configured to include multiple sub-files, and obtaining the download address of each of the sub-files stored in the remote file center, downloading each of the sub-files as a second temporary file to the local, deserializing each of the second temporary files into a sub-target object in memory, and aggregating each of the sub-target objects to form the final target object.
[0018] Optionally, in the data processing method according to one aspect of the embodiment, when receiving a signal that the obtained data is incorrect, querying the remote file center for the historical version of the file related to the data, and obtaining data based on the rolled-back historical version of the file.
[0019] To achieve the above object, according to a second aspect of the embodiments of the present invention, there is provided a data processing apparatus, including:
[0020] A data production unit: pulling raw data from a data source and performing processing;
[0021] A data writing unit: receiving the intermediate data obtained through the processing and writing the intermediate data to the remote file center, and sending a notification message with a predetermined content related to the intermediate data to the message processing center, wherein the change of the predetermined content of the notification message can be monitored by the client.
[0022] To achieve the above object, according to a second aspect of the embodiments of the present invention, there is also provided a data processing apparatus, including:
[0023] A monitoring unit: monitoring the notification message with a predetermined content sent by the server to the message processing center, monitoring the change of the predetermined content of the notification message, wherein the predetermined content of the notification message is related to the intermediate data obtained through processing the raw data; and
[0024] A data acquisition unit: in response to the change of the predetermined content of the notification message, receiving callback information from the message processing center, and parsing the message information of the callback information to obtain data.
[0025] To achieve the above object, according to a third aspect of an embodiment of the present invention, an electronic device for data processing is provided, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in one aspect of the embodiment.
[0026] To achieve the above object, according to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, characterized in that when the program is executed by a processor, the method described in one aspect of the embodiment is implemented.
[0027] One embodiment of the above invention has the following advantages or beneficial effects: Since data processing adopts the technical means of pulling raw data from a data source and processing it; receiving intermediate data obtained through the processing and writing the intermediate data into a remote file center, and sending a notification message with a predetermined content related to the intermediate data to a message processing center, the offline data acquisition and loading are independent into an independent service. Thus, the service alone undertakes the function of offline data loading and acquisition, making the offline and online completely decoupled. Additionally, since data processing adopts the technical means of listening to a notification message with a predetermined content sent from a server to a message processing center, monitoring changes in the predetermined content of the notification message, where the predetermined content of the notification message is related to intermediate data obtained through processing of raw data; and in response to changes in the predetermined content of the notification message, receiving callback information from the message processing center and parsing the message information of the callback information to obtain data, therefore, this online service can be independent of the offline service, completely decoupled from the offline service, improving the stability of the online service, reducing resource consumption, reducing resource contention for the online service, and improving the stability of the online service. The online service is decoupled from the offline data source and does not directly depend on the offline data source, and the architecture split can be fully carried out; reducing initialization logic: when the service starts, only the data file needs to be loaded, reducing unnecessary network I / O (input and output of data) and data aggregation processing, and all operations such as parsing and processing of data are completely borne by the offline data service; reducing the pressure on the data source: the offline data will be completely obtained by the offline data service at regular intervals, reducing the access pressure on the data source (Mysql, Redis, Rpc, http).
[0028] Moreover, other technical means in one aspect of an embodiment of the above invention can also achieve the following: improvement in version control and fault tolerance ability, that is, promoting multi-version data, enabling quick rollback to the previous version in case of exceptions, and providing manual intervention measures and a perfect monitoring environment; and reduction of unnecessary data replacement, that is, pulling data for online services through listening, so as to ensure that the data loaded each time must be the data with changes, rather than obtaining the same data as in the prior art, thus reducing resource waste.
[0029] The further effects of the above non-conventional optional means will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings are used to better understand the present invention and do not unduly limit the present invention. Among them:
[0031] Figure 1 is a schematic diagram of the main process of the data processing method according to an embodiment of the present invention;
[0032] Figure 2 is a schematic diagram of the main process of the data processing method according to an embodiment of the present invention;
[0033] Figure 3 is a schematic flowchart of an example of the data processing method according to an embodiment of the present invention;
[0034] Figure 4 is a schematic block diagram of an example of the data processing method according to an embodiment of the present invention;
[0035] Figure 5 is a schematic diagram of the main modules of the data processing apparatus according to an embodiment of the present invention;
[0036] Figure 6 is a schematic diagram of the main modules of the data processing apparatus according to an embodiment of the present invention;
[0037] Figure 7 is an exemplary system architecture diagram to which the embodiment of the present invention can be applied;
[0038] Figure 8 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0040] Before specific description, the following explanations are made for the terms that may appear in the embodiments or the drawings of the specification, so as to more clearly understand the technical solutions.
[0041] Mysql: Mysql is a relational database management system. Relational databases store data in different tables instead of putting all data in a large warehouse, which increases speed and flexibility.
[0042] RPC: RPC is the abbreviation of Remote Procedure Call. There are some C / S systems similar to the three-tier architecture. Third-party client programs call internal standard or custom functions through interfaces, obtain the data returned by the functions, and display or print after processing.
[0043] Redis (Remote Dictionary Server), that is, remote dictionary service, is an open-source log-type, Key-Value database written in ANSI C language, supporting network, and can be based on memory or persistent.
[0044] HDFS: An open-source distributed file storage system.
[0045] CPU: The Central Processing Unit (CPU), as the operation and control core of a computer system, is the final execution unit for information processing and program operation.
[0046] GC: Garbage Collection. In the computer field, it means that when the dynamic memory (memory space) on a computer is no longer needed, it should be released to make room for other uses. This kind of memory resource management is called garbage collection.
[0047] Network IO: IO refers to the input and output of data. To put it simply, it is the process of data moving from one place to another in a computer. In a computer, the medium that can store data is the storage medium, and the common storage structures are arrays, linked lists, and trees. We know that data in the network is transmitted in binary form. Therefore, it can be simply considered that data is the elements in a binary array. Network IO is actually the process of data moving from one array to another.
[0048] OSS: The full name of OSS is Object Storage Service, that is, object storage service.
[0049] JSON: JSON (JavaScript Object Notation) is a lightweight data interchange format. It is easy for humans to read and write. At the same time, it is also easy for machines to parse and generate.
[0050] Serialization and deserialization: Serialization is the process of converting the state information of an object into a form that can be stored or transmitted, that is, the process of converting an object into a byte sequence is called object serialization.
[0051] Deserialization is the reverse process of serialization. Deserializing a byte array into an object and restoring the byte sequence to an object is called object deserialization.
[0052] Debug: The process of debugging a program. Breakpoints are set in the program, and the program data at each step is output for debugging.
[0053] Figure 1 It is a schematic diagram of the main process of the data processing method according to the embodiments of the present invention. As Figure 1 shown, it includes step S101 to step S102. These step S101 to step S102 correspond to the offline processing of data, for example, executed by the server, such as the server write component.
[0054] Step S101: Pull the original data and perform processing.
[0055] This step S101 is part of the data production stage. It obtains data from various data sources (such as Mysql, Redis, Rpc, http, etc.); loads the original data into memory for efficient calculation and query operations; according to the requirements of the business party, performs aggregation query operations on the data, such as summation, average value, counting, etc. and / or filtering query operations to filter out data that meets specific conditions or other data calculation logics. During the entire data processing process, the data is processed in memory, so the processing speed is relatively fast.
[0056] According to step S101, processed data can be obtained in memory. For the convenience of this embodiment, it is called intermediate data, and it can be exported to a file, database, or other target locations as needed for subsequent use. Step S101 can be a one-time data processing task to meet the requirements of the business party, or it can be a part of a regularly executed data processing job.
[0057] Step S102: Receive the intermediate data obtained through the processing and write the intermediate data to the remote file center, and send a notification message with predetermined content related to the intermediate data to the message processing center. Among them, the change of the predetermined content of the notification message can be monitored by the client.
[0058] This step S102 is equivalent to the data export process on the server side, which can export and store the data produced in step S101 in the remote file center (for example, HDFS or OSS) and send a notification message.
[0059] More specifically, first, the full amount of the processed intermediate data in step S101 is serialized into Json / pb format (non-limiting, only an example for illustration, and the serialization protocol can be customized, such as Kryo, Protobuf, Hessian, etc.). Then, these serialized data are stored in a local temporary file. Next, through the API of HDFS or OSS, this temporary file is uploaded to HDFS or OSS, and then the local temporary file is deleted. Through this operation, the processed data is stored in the distributed storage system for subsequent use. Finally, a notification message is sent to consul, that is, the message processing center.
[0060] The content of the notification message includes, for example, the following items (non-limiting, only examples, and can be set as needed):
[0061] The unique identifier of the business data (the element used to uniquely identify the business data)
[0062] Generation time (the timestamp when the data is generated)
[0063] The number of generated business data (the quantity of the generated business data)
[0064] The oss address of the file (the address of the file storing the data in OSS)
[0065] Serialization type (JSON / PB) (the element indicating the data serialization type)
[0066] Business custom extension fields (some extension fields customized by the business party, which can be used for subsequent data processing or business expansion)
[0067] Through the above process, the data generation phase is completed.
[0068] As described above, in this embodiment, in steps S101 and S102, offline tasks are performed. Data is pulled from the original data source, and operations such as data calculation and aggregation query are performed according to the requirements of the business party. Finally, the calculation results are serialized into files in Json or Pb format and stored in local temporary files. Then, these files are uploaded to the distributed storage system through the API of HDFS or OSS, and the local temporary files are deleted. In this stage, all data processing and calculations are completed locally, so it is the offline stage.
[0069] Figure 2 It is a schematic diagram of the main process of the data processing method according to the embodiment of the present invention. As Figure 2 shown, it includes steps S201 to S202, and these steps S201 to S202 correspond to the online processing of data. For example, it is executed by the client, such as the client read component.
[0070] Step S201: Listen to a notification message with a predetermined content sent by the server to the message processing center, and monitor changes in the predetermined content of the notification message. Among them, the predetermined content of the notification message is related to the intermediate data obtained by processing the original data.
[0071] Step S202: In response to the change in the predetermined content of the notification message, receive callback information from the message processing center, and parse the message information of the callback information to obtain data.
[0072] In steps S201 and S202, real-time data processing and verification can be performed after receiving a new notification message, which corresponds to the online processing of data. Therefore, the above steps S201 - S202 can separate the online processing stage and the offline processing stage of data.
[0073] The offline stage mainly conducts large-scale computing and data processing locally, while the online stage performs real-time data processing and verification after receiving new notification messages. This separation design can improve data processing efficiency and reduce the impact on online data users. The offline and online are completely decoupled, enhancing the stability of online services, reducing resource consumption, minimizing resource contention for online services, and improving the stability of online services; the online service is decoupled from the offline data source and does not directly rely on the offline data source, enabling complete architecture splitting; reducing initialization logic: when the service starts, it only needs to load data files, reducing unnecessary network I / O and data aggregation processing. All operations such as data parsing and processing are fully borne by the offline data service; alleviating the pressure on the data source: the offline data is completely fetched by the offline data service at regular intervals, reducing the access pressure on the data source (Mysql, Redis, Rpc, http).
[0074] In an example of this embodiment, regarding step S201, the operations of the client may include: data pulling: when there is any change in consul (message processing center), it will send a callback message to the listener of this path, that is, the client. After the client monitors this message, it performs subsequent file processing; data construction: the client parses the information in the callback message, obtains the URL address for file download after getting it, downloads the file to a temporary file locally, and then the client deserializes this file, converts it into a Java object in memory (non-limiting, it can be other target objects), and deletes this temporary file locally; data verification: data verification is divided into basic data verification and / or custom data verification. Basic data verification compares the number of data items in the callback message with the number of data items in the deserialized Java object. If there is a difference between the two, the verification will fail and an alarm will be issued. Custom data verification is based on the rules customized by the business side, providing a data verification rule interface to support the business side to customize rules. The input of the verification rule interface is the deserialized Java object and the data in the callback message, and the output of the verification rule interface is true or false; and data switching: the new Java object replaces the old object in memory (in this scenario, the data does not allow the data listener to process and modify it, so there is no abnormal situation of overwriting and changing the old object in memory). The entire data processing, verification, and switching process is imperceptible to the data user.
[0075] Figure 3 and Figure 4 The flowchart and block diagram of an example are shown when considering the offline processing and online processing as a system, which schematically shows the main process of the offline and online data processing method of this embodiment.
[0076] In step S301, the original data is pulled from the data source (the database in this example) to obtain the business data (data acquisition); in step S302, the data calculation logic such as the business-side aggregation query and filtering query operations is completed according to the requirements of the business party (data generation); in step S303, the full data result is serialized into the Json / pb format (file format generation); in step S304, a temporary file is generated locally; in step S305, the temporary file is uploaded through the API of HDFS / OSS, stored in HDFS / OSS, the local temporary file is deleted, and a notification message of data generation is sent to consul (message processing center) (data export).
[0077] In the above steps S301 to S305, the data generation stage is completed. In this stage, all data processing and calculations are completed locally, so it is an offline stage
[0078] In step S401, when there is any change in consul (message processing center), a callback message will be sent to the listener of this path, that is, the client. After the client listens to the callback message, subsequent file processing is performed (listening to the callback message and pulling data); in step S402, the callback message information is parsed. After obtaining the file download url address, the file is downloaded to a local temporary file, and then the in-memory java object is deserialized, and the local temporary file is deleted (data construction); in step S403, data verification is performed. The data verification is divided into two parts: basic data verification and custom data verification; in the basic data verification, the number of data in the message is compared with the number of data in the deserialized object. If there is a difference, the verification fails and an alarm is issued; in the custom data verification rule, a data verification rule interface is provided to support the business party to customize the rule. The input of the verification interface is: the deserialized java object and the message data. The output of the verification interface is true or false (data verification); in step S404, the new java object replaces the old object in the memory (data switching).
[0079] According to steps S401 to 404, real-time data processing and verification can be performed after receiving a new notification message, which can be regarded as an online service and can be separated from the foregoing offline tasks.
[0080] For the data processing method of this embodiment, a safeguard mechanism can also be set.
[0081] Regarding the guarantee mechanism, the first aspect may involve scalability. For example, configure multiple machines to configure the server in a way of distributed lock and horizontal scalability, and the multiple machines are redundantly deployed and use lock contention; for another example, the original data is split into multiple data blocks, and each of the machines of the server processes one data block, generates a corresponding sub-file and stores it in the corresponding remote file center, and the client obtains the download address of each sub-file, downloads each sub-file as a second temporary file to the local, deserializes each second temporary file into a sub-target object in memory, and aggregates each sub-target object to become the final target object.
[0082] Use distributed locks to support horizontal scalability and ensure that only one machine will execute the same task. Multiple machines are redundantly deployed and lock contention is used to prevent the periodic task from not being able to execute if a certain machine fails. Moreover, since the data volume of the business side is unknown and may be hundreds of millions of records, for example, the server cannot write all the data into the same file. At this time, data splitting and sharding processing are performed. Each machine only processes a part of the data and generates a file. After the client parses and deserializes multiple files, they are aggregated into a final file.
[0083] Regarding the guarantee mechanism, the second aspect may involve monitoring and operation and maintenance tools. For example, set timeout and monitor the pulling duration of the original data, the processing duration of the original data, and the writing duration, and set timeout alarms. In addition, all file paths can be saved in the remote file center. When there is an abnormal periodic task, use the interface to query the historical version of the file and perform data version rollback.
[0084] In terms of monitoring: During the data production process, including the duration of obtaining and querying data from the data source, the duration of writing to the file, etc., set timeout and monitor, and set timeout alarms. In terms of operation and maintenance: All file paths are saved on HDFS. When a certain periodic task is abnormal and may affect the online situation, the interface can be used to query the historical version and perform data version rollback; provide a debug interface to debug the data processing situation of each process node of the data processing, analyze the data situation and abnormal situation for data joint debugging and problem troubleshooting.
[0085] Regarding the guarantee mechanism, the third aspect may involve data verification: file MD5 verification, data count verification, and custom verification (specifically, see the data verification part of the server mentioned above).
[0086] Above, the method for data processing in the embodiments of the present invention has been described with reference to the accompanying drawings.
[0087] In terms of business practice, in traditional methods, the system (such as the settlement system) does not split offline tasks and online services, resulting in a soaring database CPU. Due to reasons such as the complexity of the system in settlement, there are many offline jobs executed, such as regularly verifying the accuracy of doctors' account periods and regularly splitting the settlement data amortized by day. Since these data rely on the database for execution, the CPU of the main settlement database is very high, usually stable above 30% (the CPU usage of other projects does not exceed 10%). In case of large promotional activities such as 618 with high traffic, offline tasks need to be stopped to ensure the stability of online interfaces. The consequence of suspending offline tasks is that, for example, for the task of timely amortizing cost data, it is difficult to ensure that the switch can be turned on manually in time later. If the amortized data is incomplete, it will affect financial reconciliation and the accuracy of daily, weekly, and monthly data.
[0088] In contrast, according to the data processing method of this embodiment, the system (such as the settlement system) splits offline tasks and online services. After the split settlement system, the database CPU usage remains below 10%, and offline scheduled tasks do not need to be suspended during large promotional activities, ensuring the accuracy and real-time nature of settlement data.
[0089] An example of the above system applying the data processing method of this embodiment is the settlement system, but this is not restrictive and can be applied to any system.
[0090] Using the data processing method of this embodiment, the following are achieved: decoupling of offline and online: reducing resource consumption: during the execution of scheduled tasks, there is competition for system resources with online services. After task splitting, resource competition is reduced; reducing the dependency relationship of scheduled tasks, reducing the blocking of scheduled tasks, and optimizing the problem of long service startup time. Improvement of data fault tolerance: providing version fault tolerance function: multiple version control of data files, and abnormal data sources can be rolled back to the previous version; flexible operation and maintenance tools and means: better ability of manual intervention to ensure quick recovery of problems; version control ability: avoiding frequent online loading of duplicate data. Improvement of system stability: avoiding the direct dependence of online services on offline data sources. Reduction of offline pressure: reducing the access pressure on data sources such as databases and reducing the access frequency. Reducing data coupling.
[0091] Figure 5 Fig. 500 shows a data processing apparatus provided in the second aspect of the embodiment of the present invention, including: a data production unit 501: pulling raw data from a data source and performing processing; a data writing unit 502: receiving the intermediate data obtained through the processing and writing the intermediate data into a remote file center, and sending a notification message with predetermined content related to the intermediate data to a message processing center. Wherein, the change of the predetermined content of the notification message can be monitored by a client.
[0092] The data processing device 500 corresponds to the device for processing offline data. It can be set up on the server side and can achieve the same technical effects as the above-mentioned data processing method.
[0093] Figure 6 Figure 600 of the data processing device provided by the second aspect of the embodiment of the present invention is shown, including: a monitoring unit 601: monitoring a notification message with a predetermined content sent from the server to the message processing center, and monitoring changes in the predetermined content of the notification message, where the predetermined content of the notification message is related to intermediate data obtained by processing original data; and a data acquisition unit 602: in response to a change in the predetermined content of the notification message, receiving callback information from the message processing center, and parsing the message information of the callback information to obtain data.
[0094] The data processing device 600 corresponds to the device for processing online data. It can be set up on the client side and can achieve the same technical effects as the above-mentioned data processing method.
[0095] Because the data processing adopts the technical means of pulling original data from the data source and processing it; the server receives the processed intermediate data and writes the intermediate data to the remote file center, and sends a notification message with a predetermined content; and the client responds to a change in the predetermined content of the notification message, receives callback information, and parses the message information of the callback information to obtain data and perform data verification and use, so it overcomes the technical problem of offline and online coupling, and thus achieves the following technical effects: offline and online are completely decoupled, improving the stability of online services, reducing resource consumption, reducing the competition for online service resources, and improving the stability of online services; online services are decoupled from offline data sources, do not directly depend on offline data sources, and the architecture can be completely split; reduce initialization logic: only need to load data files when the service starts, reducing unnecessary network I / O and data aggregation processing, and all operations such as parsing and processing of data are completely carried by the offline data service; relieve the pressure on the data source: offline data will be completely obtained by the offline data service at regular intervals, reducing the access pressure on the data source (Mysql, Redis, Rpc, http).
[0096] Moreover, version control and fault tolerance are improved: multiple versions of data are improved, exceptions can be quickly rolled back to the previous version, and manual intervention measures and a perfect monitoring environment are provided; and unnecessary data replacement is reduced: online services pull data through monitoring, so it can be ensured that the data loaded each time must be the data with changes, rather than obtaining the same data as in the old technology, reducing waste of resources.
[0097] Figure 7FIG. 700 shows an exemplary system architecture to which the data processing method or data processing apparatus according to an embodiment of the present invention can be applied.
[0098] As Figure 7 shown, the system architecture 700 may include terminal devices 701, 702, 703, a network 704, and a server 705 (this architecture is only an example, and the units included in a specific architecture can be adjusted according to the specific situation of the application). The network 704 is used to provide a medium for communication links between the terminal devices 701, 702, 703 and the server 705. The network 704 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0099] Users can use the terminal devices 701, 702, 703 to interact with the server 705 through the network 704 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 701, 702, 703, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only examples).
[0100] The terminal devices 701, 702, 703 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0101] The server 705 may be a server providing various services, such as a background management server that supports shopping websites browsed by users using the terminal devices 701, 702, 703 (only an example). The background management server may analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - only examples) to the terminal devices.
[0102] It should be noted that the data processing method provided by the embodiments of the present invention is generally executed by the server 705. Correspondingly, the data processing apparatus is generally disposed in the server 705.
[0103] It should be understood that Figure 7 the numbers of the terminal devices, the network, and the server in
[0104] are only illustrative. According to actual needs, there may be any number of terminal devices, networks, and servers. Figure 8 are only illustrative. According to actual needs, there may be any number of terminal devices, networks, and servers. Figure 8 The terminal device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0105] As Figure 8As shown, computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0106] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0107] Specifically, according to an embodiment disclosed by the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment disclosed by the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a central processing unit (CPU) 801, the above functions defined in the system of the present invention are executed.
[0108] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0110] The units (or "modules") involved in the embodiments of the present invention can be implemented in software or in hardware. The described units (or "modules") can also be provided in a processor. For example, it can be described as: a processor includes a data production unit (or "module") and a data writing unit; for example, it can be described as a processor includes a monitoring unit (or "module") and a data acquisition unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the data production unit can also be described as "a unit that pulls raw data and processes it".
[0111] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; it can also exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device includes: pulling raw data and processing it; and receiving intermediate data obtained through the processing and writing the intermediate data into a remote file center, and sending a notification message with a predetermined content related to the intermediate data to a message processing center, where the change of the predetermined content of the notification message can be monitored by a client. It can also be that the above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device includes: monitoring a notification message with a predetermined content sent by a server to a message processing center, monitoring the change of the predetermined content of the notification message, where the predetermined content of the notification message is related to intermediate data obtained through processing of raw data; and in response to the change of the predetermined content of the notification message, receiving callback information from the message processing center and parsing the message information of the callback information to obtain data.
[0112] According to the technical solution of the embodiments of the present invention, it is possible to separate offline data from online services, relieve the pressure on the data source, make the data production process flow-based and configurable, and improve the fault tolerance ability. The acquisition and loading of offline data are independent into an independent service, and this service alone undertakes the function of loading and acquiring offline data.
[0113] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for data processing, characterized in that: include: Pull the raw data and process it; as well as Receive the intermediate data obtained through the processing and write the intermediate data to a remote file center, and send a notification message with predetermined content related to the intermediate data to a message processing center, wherein changes in the predetermined content of the notification message can be monitored by the client.
2. The data processing method according to claim 1, characterized in that: The processing of the raw data includes: In response to user needs, the original data of the server is pulled and stored in the memory, and the stored original data is processed in the memory to generate the intermediate data.
3. The data processing method according to claim 2, characterized in that: The receiving the intermediate data obtained through the processing and writing the intermediate data to a remote file center includes: Serializing the full amount of the intermediate data into a predetermined format and generating a first temporary file locally; and The generated first temporary file is stored in the remote file center, and the first temporary file locally is deleted.
4. The data processing method according to claim 3, characterized in that: The method is implemented by configuring multiple machines in a distributed lock and horizontally scalable manner. The multiple machines are redundantly deployed and use contention locks, and The original data is split into multiple data blocks, and each of the machines processes a data block to generate a corresponding sub-file and store it in a remote file center.
5. The data processing method according to claim 4, characterized in that: Also includes: The time duration for pulling the original data, the time duration for processing the original data, and the time duration for writing are set and monitored overtime, and a timeout alarm is set.
6. A method of data processing, characterized in that: include: Monitoring a notification message with predetermined content sent by a server to a message processing center, and monitoring a change of the predetermined content of the notification message, wherein the predetermined content of the notification message is related to intermediate data obtained by processing the original data; as well as In response to the change of the predetermined content of the notification message, callback information is received from a message processing center, and message information of the callback information is parsed to obtain data.
7. The data processing method according to claim 6, characterized in that: The parsing of the message information of the callback information to obtain data includes: Based on the predetermined content of the notification message, obtain a file download address of a file stored in a remote file center, download the file locally as a second temporary file, deserialize the second temporary file into a target object in memory, delete the local second temporary file, and verify the target object and data in the callback message, and When the verification passes, it triggers the replacement of the old target object in the memory with the target object.
8. The data processing method according to claim 7, characterized in that: The file is configured to include a plurality of sub-files, and Obtain the download address of each of the sub-files stored in the remote file center, download each of the sub-files to the local as a second temporary file, deserialize each of the second temporary files into a sub-target object in the memory, and aggregate each of the sub-target objects to become the final target object.
9. The data processing method according to claim 7 or 8, characterized in that: When a signal indicating that the acquired data has errors is received, the remote file center is queried for historical versions of files related to the data, and data is acquired based on the rolled-back historical versions of the files.
10. A data processing device, characterized in that: include: Data production unit: pulls raw data and processes it; as well as Data writing unit: receives the intermediate data obtained through the processing and writes the intermediate data to a remote file center, and sends a notification message with predetermined content related to the intermediate data to a message processing center, wherein changes in the predetermined content of the notification message can be monitored by the client.
11. A data processing device, characterized in that: include: A monitoring unit: monitors a notification message with predetermined content sent by a server to a message processing center, and monitors changes in the predetermined content of the notification message, wherein the predetermined content of the notification message is related to intermediate data obtained by processing the original data; as well as The data acquisition unit receives callback information from a message processing center in response to a change in the predetermined content of the notification message, and parses the message information of the callback information to acquire data.
12. An electronic device for data processing, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.
13. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.