User data extraction method and device, electronic equipment and storage medium
By setting the selection state and calculation index during the user data extraction process, the problem of sample data extraction in the prior art cannot be reproduced, and the extraction efficiency and fairness are improved.
Patent Information
- Application Number
- CN202510459471.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
In the prior art, there is current limiting and randomness in the process of user data extraction, which leads to the inability to reproduce the sample data extraction, affecting business fairness.
By obtaining user data based on user information, determining the number to be selected, calculating the index and offset of user data, setting the selection status, and accurately extracting sample data.
It improves the efficiency, fairness and replayability of sample extraction, reduces network overhead, and enhances the compatibility and expansion of the system.
Smart Images

Figure CN119988367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technology, and in particular to a user data extraction method, device, electronic device and storage medium. Background Art
[0002] In the prior art, when extracting sample data from user data, a flow limiting method is usually used to filter out most of the user data, and sample data is randomly obtained from the remaining user data. Due to the uncertainty of flow limiting and random extraction, sample data extraction cannot be replayed, which brings risks to business fairness. Summary of the invention
[0003] In view of this, an object of the present invention is to provide a user data extraction method, device, electronic device and storage medium, which can improve the efficiency, fairness and replayability of sample extraction.
[0004] In order to achieve the above purpose, the technical solution adopted by the embodiment of the present invention is as follows: In a first aspect, the present invention provides a method for extracting user data, the method comprising: According to the user information, corresponding user data is obtained; the user information is used to characterize the users participating in the sample extraction; The number of data to be selected is determined according to the total number of data and the number of samples; the total number of data is the total number of all user data; the number of samples is the total number of sample data; Determine a step size according to the index of the user data, the offset, the total amount of data and the amount to be selected, and determine a minimum index and a maximum index according to the step size, the offset, the total amount of data and the amount to be selected; the offset represents a starting index of the user data to be selected; When the index of the user data is the minimum index or the maximum index, setting the selection state corresponding to the user data to be selected; According to the selection state, the sample data is determined from the user data.
[0005] In an optional implementation manner, determining the number to be selected according to the total number of data and the number of samples includes: When the total amount of data is greater than the sample amount, the non-sample amount is obtained according to the difference between the total amount of data and the sample amount; When the sample quantity is less than the non-sample quantity, determining the sample quantity as the quantity to be selected; When the sample quantity is not less than the non-sample quantity, the non-sample quantity is determined as the to-be-selected quantity.
[0006] In an optional implementation manner, determining the step size according to the index, offset, total amount of data and amount to be selected of the user data, and determining the minimum index and the maximum index according to the step size, offset, total amount of data and amount to be selected, comprises: Get the offset; When the index of the user data does not exceed the offset, setting the selection state corresponding to the user data to unselected; When the index of the user data exceeds the offset, determining a fixed tolerance according to the total amount of data and the amount to be selected; Determine a step size according to the index of the user data, the offset and the fixed tolerance; The minimum index and the maximum index are determined according to the step size, the fixed tolerance, and the offset.
[0007] In an optional implementation manner, determining the sample data from the user data according to the selection state includes: When the sample quantity is less than the non-sample quantity, determining the user data in the selected state as the sample data; When the sample quantity is not less than the non-sample quantity, the user data whose selection status is unselected is determined as the sample data.
[0008] In an optional implementation manner, acquiring corresponding user data according to the user information includes: Concurrently obtaining the user information from an extraction queue; the extraction queue is used to store user information corresponding to users participating in sample extraction; Determine a corresponding user queue according to the user information; the user information corresponds to the user queue one by one; the user queue is used to store user data corresponding to the user information; The user queue is traversed to obtain all user data corresponding to the user information.
[0009] In an optional implementation manner, before the step of acquiring corresponding user data according to the user information, the method further includes: Receive a sample extraction request from a user; the sample extraction request includes a user identifier and a number of extractions; Generate the user information according to the user identifier, and add the user information to the extraction queue; Generate user data according to the user identifier and the extraction times, and add the user data to a user queue.
[0010] In an optional implementation manner, generating user data according to the user identifier and the number of extractions includes: When the sum of the current total number of extractions corresponding to the user identifier and the number of extractions does not exceed the extraction threshold, obtaining the minimum index value and the maximum index value corresponding to the sample extraction request according to the number of extractions; User data corresponding to the sample extraction request is generated according to the user identifier, the minimum index value and the maximum index value.
[0011] In a second aspect, the present invention provides a user data extraction device, the device comprising: An acquisition module, used to acquire corresponding user data according to user information; the user information is used to characterize the users participating in sample extraction; A processing module, configured to determine the number to be selected according to the total number of data and the number of samples; the total number of data is the total number of all user data; the number of samples is the total number of sample data; a step length is determined according to the index of the user data, an offset, the total number of data and the number to be selected, and a minimum index and a maximum index are determined according to the step length, the offset, the total number of data and the number to be selected; the offset represents the starting index of the user data to be selected; when the index of the user data is the minimum index or the maximum index, the selection state corresponding to the user data is set to be selected; An extraction module is used to determine the sample data from the user data according to the selection state.
[0012] In a third aspect, the present invention provides an electronic device, comprising a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the user data extraction method described in any of the aforementioned implementations.
[0013] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the user data extraction method as described in any one of the aforementioned embodiments.
[0014] Compared with the prior art, the user data extraction method, device, electronic device and storage medium provided in the embodiments of the present invention obtain corresponding user data according to user information; determine the number to be selected according to the total number of data and the number of samples; determine the minimum index and the maximum index according to the index of the user data, the total number of data and the number to be selected; when the index of the user data is the minimum index or the maximum index, set the selection state corresponding to the user data to selected; and determine sample data from the user data according to the selection state.
[0015] It can be seen that the embodiment of the present invention obtains all user data of each user at one time when extracting sample data, which can greatly reduce the network overhead of obtaining user data and greatly improve the efficiency of sample extraction. By assigning a unique index to each user data and accurately extracting sample data using the index of the user data, the total number of user data and the total number of sample data, the openness, fairness and replayability of user data extraction are effectively improved, so that it can show excellent compatibility and scalability when facing higher traffic requests and shorter time requirements.
[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A schematic flow chart of a method for extracting user data provided by an embodiment of the present invention is shown.
[0019] Figure 2 Another schematic flow chart of a method for extracting user data provided by an embodiment of the present invention is shown.
[0020] Figure 3 Another schematic flow chart of a method for extracting user data provided by an embodiment of the present invention is shown.
[0021] Figure 4 A block diagram of a user data extraction device provided by an embodiment of the present invention is shown.
[0022] Figure 5 A block diagram of an electronic device provided by an embodiment of the present invention is shown.
[0023] Icon: 100 - electronic device; 110 - memory; 120 - processor; 130 - communication module; 300 - user data extraction device; 301 - acquisition module; 302 - processing module; 303 - extraction module. DETAILED DESCRIPTION
[0024] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0026] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0027] In today's Internet industry, there is a need to extract user data in situations involving massive user participation. For example, various online platforms have set up lottery activities, and the process of drawing prizes can essentially be considered as the process of extracting user data.
[0028] In the face of the impact of high traffic on the system, it is difficult to achieve absolutely fair sample extraction. In order to ensure the fairness of sample extraction, most user data is usually filtered out through flow limiting, and then a random number generation algorithm is used to determine whether the user data can be extracted as sample data. Due to the uncertainty of flow limiting and random number algorithms themselves, the user data extraction process cannot be restored on site or replayed, which brings risks to the fairness of the business.
[0029] Based on this, the user data extraction method and device provided by the embodiment of the present invention can obtain all user data of each user at one time when performing sample data extraction, which can greatly reduce the network overhead of obtaining user data and greatly improve the efficiency of sample extraction. By assigning a unique index to each user data and accurately extracting sample data using the index of the user data, the total number of user data and the total number of sample data, the openness, fairness and replayability of user data extraction are effectively improved, so that it can show excellent compatibility and scalability when facing higher traffic requests and shorter time requirements.
[0030] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0031] See also Figure 1 , Figure 1 A schematic diagram of a flow chart of a user data extraction method provided by an embodiment of the present invention is shown. The method should include the following steps: Step S40: acquiring corresponding user data according to the user information.
[0032] Among them, user information is used to characterize the users participating in sample extraction.
[0033] In the embodiment of the present invention, obtaining user data requires user consent or authorization. User information can uniquely identify the user participating in sample extraction, and the user information can include a user account, user ID, or user name. Based on the user information, all user data corresponding to the user, i.e., target user data, is obtained.
[0034] Step S50, determining the number of items to be selected based on the total number of data and the number of samples.
[0035] The total amount of data is the total amount of all user data, and the sample amount is the total amount of sample data.
[0036] In the embodiment of the present invention, the user data extraction process is determined according to the size of the total amount of data and the sample amount. When the total amount of data is less than or equal to the sample amount, all user data corresponding to the user information is determined as sample data, that is, all user data corresponding to each user is determined as sample data. For example, in the lottery stage, everyone has a prize, that is, the number of prizes is the same as the number of draws, that is, the sample amount is the same as the total amount of data.
[0037] When the total amount of data is greater than the number of samples, the number to be selected is determined according to the total amount of data and the number of samples, wherein the number to be selected represents the number of user data in the selection state being selected.
[0038] Step S60, determining the step size according to the index, offset, total amount of data and amount to be selected of the user data, and determining the minimum index and the maximum index according to the step size, offset, total amount of data and amount to be selected.
[0039] The offset represents the starting index of the user data to be selected.
[0040] Step S70, determining whether the index of the user data is a minimum index or a maximum index.
[0041] In the embodiment of the present invention, when the index of the target user data is equal to the minimum index or the maximum index, step S80 is executed. When the index of the target user data is not equal to the minimum index and the index of the target user data is not equal to the maximum index, step S90 is executed. The index of the user data may be a positive integer user data ID (e.g., 1, 100, etc.), which is globally uniformly addressed and used to uniquely identify the user data.
[0042] Step S80: If yes, set the selection status corresponding to the user data to selected.
[0043] Step S90: If not, set the selection status corresponding to the user data to unselected.
[0044] Step S100, determining sample data from user data according to the selection status.
[0045] In an embodiment of the present invention, the selection status corresponding to the target user data whose index is the minimum index or the maximum index is set to selected, and the selection status corresponding to the target user data whose index is neither the minimum index nor the maximum index is set to unselected. Finally, sample data is selected from the target user data based on the selection status, and the extraction result of the sample data is notified to the user.
[0046] In summary, the user data extraction method provided in the embodiment of the present invention obtains corresponding user data according to user information; determines the number of data to be selected according to the total number of data and the number of samples; determines the minimum index and the maximum index according to the index of the user data, the total number of data and the number of data to be selected; when the index of the user data is the minimum index or the maximum index, sets the selection state corresponding to the user data to selected; and determines the sample data from the user data according to the selection state.
[0047] It can be seen that the embodiment of the present invention obtains all user data of each user at one time when extracting sample data, which can greatly reduce the network overhead of obtaining user data and greatly improve the efficiency of sample extraction. By assigning a unique index to each user data and accurately extracting sample data using the index of the user data, the total number of user data and the total number of sample data, the openness, fairness and replayability of user data extraction are effectively improved, so that it can show excellent compatibility and scalability when facing higher traffic requests and shorter time requirements.
[0048] Optionally, when the total number of data is greater than the number of samples, the following is a possible implementation method for determining the number of samples to be selected. Figure 2 The sub-steps of step S50 may include: Step S501, when the total number of data is greater than the number of samples, the non-sample number is obtained according to the difference between the total number of data and the number of samples.
[0049] Step S502: when the sample quantity is less than the non-sample quantity, the sample quantity is determined as the quantity to be selected.
[0050] Step S503: when the sample quantity is not less than the non-sample quantity, the non-sample quantity is determined as the quantity to be selected.
[0051] In the embodiment of the present invention, assuming that the total amount of data is greater than the sample amount, but less than or equal to twice the sample amount, the non-sample amount is determined according to the difference between the total amount of data and the sample amount, and the sample amount is greater than or equal to the non-sample amount. In order to improve the extraction efficiency of user data, non-sample data is selected from the user data, and the non-sample amount is used as the amount to be selected. The non-sample amount is the amount of user data that has not become sample data, and the non-sample data is the user data that has not become sample data.
[0052] Assume that the total amount of data is greater than twice the number of samples, and the number of samples is less than the number of non-samples. In order to improve the efficiency of extracting user data, sample data is selected from the user data, and the number of samples is used as the number to be selected.
[0053] It can be seen that the embodiment of the present invention determines the user data to be selected based on the size of the sample number and the non-sample number. When the sample number is less than the non-sample number, the sample data is selected from the user data. When the sample number is not less than the non-sample number, the non-sample data is selected from the user data. This ensures that a smaller number of data to be selected is obtained, effectively improving the efficiency of extracting user data.
[0054] Optionally, the following provides a possible implementation method for determining the minimum index and maximum index corresponding to each user data. Figure 2The sub-steps of step S60 may include: Step S601, obtaining an offset.
[0055] In the embodiment of the present invention, it is assumed that the system starts a single thread to count the number of samples and the total number of data. In order to ensure the fairness of sample data extraction, an offset is randomly generated according to the total number of data and the number of data to be selected, and the offset is recorded to ensure the replay of the sample data extraction process. The calculation formula of the offset is as follows: offset=rand(floor(total / selector)) Among them, offset is the offset; total is the total number of data; selector is the number to be selected; floor is the floor rounding function; rand is the random function.
[0056] Step S602: When the index of the user data does not exceed the offset, the selection state corresponding to the user data is set to unselected.
[0057] In the embodiment of the present invention, the offset is used as the starting index for extracting the user data to be selected. The user data with an index less than the offset is not selected, and the selection status corresponding to the user data is set to unselected. For example, the total number of data is 100, the indexes of all user data are from 1 to 100, and the number of selectors to be selected is 6, then the offset is a random value not exceeding 16, assuming that the offset is 5. Then when the index of the user data is less than or equal to 5, the user data is not selected. The user data with the selection status of unselected will skip steps S603-step S605, and steps S70-step S90, and then directly execute step S100.
[0058] Step S603: When the index of the user data exceeds the offset, a fixed tolerance is determined according to the total amount of data and the amount to be selected.
[0059] In the embodiment of the present invention, the ratio of the total number of data to the number to be selected is determined as a fixed tolerance. The calculation formula of the fixed tolerance is as follows: tolerance=total / selector Among them, tolerance is a fixed tolerance. For example, if the total number of data is 100 and the number of selectors to be selected is 6, then the fixed tolerance is a decimal greater than 16 and less than 17. If one decimal place is retained, the tolerance is 16.7.
[0060] Step S604, determining the step length according to the index, offset and fixed tolerance of the user data.
[0061] Step S605, determining the minimum index and the maximum index according to the step size, the fixed tolerance and the offset.
[0062] In the embodiment of the present invention, the step length, the minimum index and the maximum index are calculated for each user data. Specifically, the difference between the index and the offset of the user data is calculated to obtain the deviation value, and the step length is determined according to the ratio of the deviation value to the fixed tolerance.
[0063] The calculation formulas for the step size, minimum index, and maximum index are as follows: step=floor((index-offset) / tolerance) index_min=round(step×tolerance+offset+1) index_max=round((step+1)×tolerance+offset+1) Among them, step is the step size; index is the index of the user data; index_min is the minimum index; index_max is the maximum index; round is the rounding function.
[0064] It can be seen that the embodiment of the present invention can quickly determine the minimum index and the maximum index according to the fixed tolerance, use the minimum index and the maximum index to determine whether the user data is extracted as sample data, and reversely index to the user corresponding to the user data, thereby achieving fairness in sample extraction.
[0065] Optionally, a possible implementation method for selecting sample data is provided below. Figure 1 The sub-steps of step S100 may include: When the sample number is less than the non-sample number, the user data whose selection status is selected is determined as the sample data; when the sample number is not less than the non-sample number, the user data whose selection status is unselected is determined as the sample data.
[0066] In the embodiment of the present invention, when the number of samples is less than the number of non-samples, in order to improve the extraction efficiency of user data, sample data is selected from the user data, and the selection state corresponding to the selected user data is set to selected. That is, the selected user data is the sample data.
[0067] When the sample data is not less than the number of non-sample data, in order to improve the efficiency of user data extraction, non-sample data is selected from the user data, and the selection status corresponding to the selected user data is set to selected. In other words, the selected user data is the non-sample data, and the unselected user data is the sample data.
[0068] It can be seen that when extracting user data, the embodiment of the present invention extracts sample data or non-sample data according to the side with smaller data volume, thereby effectively improving the efficiency of extracting user data.
[0069] Optionally, the following provides a possible implementation method for obtaining user data of each user. Figure 3 , Figure 1 The sub-steps of step S40 may include: Step S401, concurrently obtain user information from the extraction queue.
[0070] Among them, the extraction queue is used to store user information corresponding to users participating in sample extraction.
[0071] In the embodiment of the present invention, it is assumed that the system starts multiple threads and concurrently takes out user information from the extraction queue, and each thread processes the obtained user information.
[0072] Step S402: determine the corresponding user queue according to the user information.
[0073] There is a one-to-one correspondence between user information and user queues, and the user queues are used to store user data corresponding to the user information.
[0074] In the embodiment of the present invention, each user is provided with a user queue, and the user queue is used to store all user data of the user. That is to say, there is a one-to-one correspondence between the user, the user information and the user queue, and the user queue corresponding to the user, i.e., the target user queue, can be found through the user information.
[0075] Step S403, traverse the user queue to obtain all user data corresponding to the user information.
[0076] In the embodiment of the present invention, the target user queue is traversed to extract all user data to obtain user data corresponding to the user information, namely the target user data. Sample data is extracted based on the target user data until the extraction queue is empty, completing the sample data extraction operation.
[0077] It is worth mentioning that, assuming that there are m users participating in sample extraction, and the extraction threshold of each user is n, the existing sample extraction usually obtains one extraction request from one user at a time for processing, so the sample extraction needs to be processed The sample extraction complexity is The sample extraction complexity of the embodiment of the present invention is only related to the number of users. Each time all the user data of the same user is obtained for processing, m times of processing are required to complete the sample extraction, and the sample extraction complexity is O(m).
[0078] It can be seen that the embodiment of the present invention stores the data users corresponding to multiple sample extraction requests of the same user into the corresponding user queue through a merging mechanism, and utilizes multi-threaded concurrency to perform sample extraction on all user data in each user queue, thereby greatly improving the concurrency efficiency and thus improving the user data extraction efficiency.
[0079] Optionally, the following provides a possible implementation method for generating user information and user data based on the user's sample extraction request. Figure 3 ,exist Figure 1 Before step S40, the following steps may also be included: Step S10: receiving a sample extraction request from a user.
[0080] The sample extraction request includes a user identifier and extraction times.
[0081] In an embodiment of the present invention, a user specifies the number of draws through a client to participate in a sample draw activity. The client responds to the user operation, generates a sample draw request according to the user identifier and the number of draws, and sends the sample draw request to the electronic device. The user identifier may be a user account, user ID, or user name, which is not limited by the present invention.
[0082] Step S20, generating user information according to the user identifier, and adding the user information to the extraction queue.
[0083] Step S30, generating user data according to the user identifier and the number of extractions, and adding the user data to the user queue.
[0084] In an embodiment of the present invention, user information may be added to an extraction queue in sequence according to the order in which sample extraction requests are received, a corresponding user queue may be obtained according to a user identifier, and newly generated user data may be saved in sequence to the corresponding user queue.
[0085] As a possible implementation method, the extraction queue is implemented based on the redis set, and the automatic deduplication function of the redis set is used to automatically deduplicate the user information in the extraction queue to ensure that each user information in the extraction queue is unique, thereby avoiding repeated extraction of user data and effectively improving the accuracy, fairness and efficiency of sample data extraction.
[0086] It should be noted that, when the extraction queue does not have an automatic deduplication function, it is necessary to first determine whether the user information already exists in the extraction queue. If it does, there is no need to generate new user information; if it does not exist, the user information generated according to the user identifier is added to the extraction queue to ensure the uniqueness of the user information in the extraction queue. The present invention does not limit the implementation method of ensuring the uniqueness of the user information in the extraction queue.
[0087] Optionally, a possible implementation method for generating user data is provided below. Figure 3 The sub-step of generating user data in step S30 may include: When the sum of the current total number of extractions corresponding to the user ID and the number of extractions does not exceed the extraction threshold, the minimum index value and the maximum index value corresponding to the sample extraction request are obtained according to the number of extractions; based on the user ID, the minimum index value and the maximum index value, the user data corresponding to the sample extraction request is generated.
[0088] In the embodiment of the present invention, when the user data extraction system is initialized, the configuration items are loaded to determine the number of samples and the extraction threshold. The extraction threshold is used to represent the maximum number of times a user can participate in sample extraction. In order to enhance the diversity of sample extraction, the continuous extraction function is supported. Each time a user initiates a sample extraction request, the number of extractions for this request can be specified.
[0089] After receiving the user's sample extraction request, for each sample extraction request, first obtain the user's current total extraction number based on the user ID. When the sum of the current total extraction number corresponding to the user ID and the number of extractions exceeds the extraction threshold, notify the user that the number of times he has participated in sample extraction has reached the maximum number of times, so as to avoid the user from initiating sample extraction requests without limit, which brings the security risk of the user data extraction system being flushed, and effectively improves the reliability of user data extraction.
[0090] For sample extraction requests that do not exceed the extraction threshold, the extraction times are input into the extraction counter, and the extraction counter is used to generate the maximum index value corresponding to this sample extraction request. Specifically, the extraction counter is used to automatically increase and allocate continuous indexes. The extraction counter continuously allocates the extraction times indexes starting from the last allocated index, and uses the smallest newly allocated index as the minimum index value, and the largest newly allocated index as the maximum index value.
[0091] From the minimum index value to the maximum index value, user data corresponding to each index is generated according to each newly allocated index and user identifier, wherein the user data includes the index.
[0092] It should be noted that in order to reduce the storage space occupied by user data, the index can be directly used as user data, and a user extraction record can be created according to the sample extraction request. The user extraction record includes the user identifier, the minimum index value, and the maximum index value. The user extraction record is obtained according to the user information, and the sample data is determined by using each index in the user extraction record that is greater than or equal to the minimum index value and less than or equal to the maximum index value.
[0093] As a possible implementation method, a clustered redis high-performance cache can be used to store key data such as extraction queues, user queues, and extraction thresholds, thereby supporting millions of participating users.
[0094] Based on the same inventive concept, the embodiment of the present invention also provides a user data extraction device. Its basic principle and technical effects are the same as those of the above embodiment. For the sake of brief description, for parts not mentioned in this embodiment, reference can be made to the corresponding contents in the above embodiment.
[0095] See also Figure 4 , Figure 4 The block diagram of a user data extraction device 300 provided by an embodiment of the present invention is shown. The user data extraction device 300 comprises an acquisition module 301 , a processing module 302 and an extraction module 303 .
[0096] The acquisition module 301 is used to acquire corresponding user data according to user information; the user information is used to characterize the users participating in the sample extraction; The processing module 302 is used to determine the number to be selected according to the total number of data and the number of samples; the total number of data is the total number of all user data; the number of samples is the total number of sample data; the step length is determined according to the index, offset, total number of data and the number to be selected of the user data, and the minimum index and maximum index are determined according to the step length, offset, total number of data and the number to be selected; the offset represents the starting index of the user data to be selected; when the index of the user data is the minimum index or the maximum index, the selection state corresponding to the user data is set to be selected; The extraction module 303 is used to determine sample data from the user data according to the selection state.
[0097] In summary, the user data extraction device provided by the embodiment of the present invention can obtain all user data of each user at one time when performing sample data extraction, which can greatly reduce the network overhead of obtaining user data and greatly improve the efficiency of sample extraction. By assigning a unique index to each user data and accurately extracting sample data using the index of the user data, the total number of user data and the total number of sample data, the openness, fairness and replayability of user data extraction are effectively improved, so that it can show excellent compatibility and scalability when facing higher traffic requests and shorter time requirements.
[0098] Optionally, the processing module 302 is specifically used to obtain the non-sample number based on the difference between the total data number and the sample number when the total data number is greater than the sample number; when the sample number is less than the non-sample number, determine the sample number as the number to be selected; when the sample number is not less than the non-sample number, determine the non-sample number as the number to be selected.
[0099] Optionally, the processing module 302 is specifically used to obtain an offset; when the index of the user data does not exceed the offset, the selection state corresponding to the user data is set to unselected; when the index of the user data exceeds the offset, a fixed tolerance is determined according to the total number of data and the number to be selected; a step size is determined according to the index of the user data, the offset and the fixed tolerance; and a minimum index and a maximum index are determined according to the step size, the fixed tolerance and the offset.
[0100] Optionally, the extraction module 303 is specifically used to determine the user data in the selected state as sample data when the sample number is less than the non-sample number; when the sample number is not less than the non-sample number, determine the user data in the unselected state as sample data.
[0101] Optionally, the acquisition module 301 is specifically used to concurrently obtain user information from the extraction queue; the extraction queue is used to store user information corresponding to users participating in sample extraction; the corresponding user queue is determined according to the user information; the user information and the user queue correspond one-to-one; the user queue is used to store user data corresponding to the user information; and the user queue is traversed to obtain all user data corresponding to the user information.
[0102] Optionally, the acquisition module 301 is also used to receive a sample extraction request from a user; the sample extraction request includes a user ID and a number of extractions; user information is generated based on the user ID, and the user information is added to the extraction queue; user data is generated based on the user ID and the number of extractions, and the user data is added to the user queue.
[0103] Optionally, the acquisition module 301 is specifically used to obtain the minimum index value and the maximum index value corresponding to the sample extraction request according to the number of extractions when the sum of the current total number of extractions corresponding to the user identifier and the number of extractions does not exceed the extraction threshold; and generate user data corresponding to the sample extraction request according to the user identifier, the minimum index value and the maximum index value.
[0104] See also Figure 5 , Figure 5 A block diagram of an electronic device 100 provided in an embodiment of the present invention. The electronic device 100 may be any device having a data processing function, such as a personal computer, a laptop computer, a server, etc. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. The memory 110, the processor 120, and the communication module 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.
[0105] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0106] The processor 120 is used to read / write data or programs stored in the memory 110 and execute corresponding functions. For example, when the computer program stored in the memory 110 is executed by the processor 120, the user data extraction method disclosed in the above embodiments can be implemented.
[0107] The communication module 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through a network, and to send and receive data through the network.
[0108] It should be understood that Figure 5 The structure shown is only a schematic diagram of the structure of the electronic device 100. The electronic device 100 may also include Figure 5 More or fewer components as shown, or with Figure 5 Different configurations are shown. Figure 5 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0109] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by the processor 120, the user data extraction method disclosed in the above embodiments is implemented.
[0110] The embodiment of the present invention further provides a program product, which implements the user data extraction method disclosed in the above embodiments when executed by the processor 120.
[0111] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0112] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0113] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0114] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A user data extraction method, characterized in that: The method comprises: According to the user information, corresponding user data is obtained; the user information is used to characterize the users participating in the sample extraction; The number of data to be selected is determined according to the total number of data and the number of samples; the total number of data is the total number of all user data; the number of samples is the total number of sample data; Determine a step size according to the index of the user data, the offset, the total amount of data and the amount to be selected, and determine a minimum index and a maximum index according to the step size, the offset, the total amount of data and the amount to be selected; the offset represents a starting index of the user data to be selected; When the index of the user data is the minimum index or the maximum index, setting the selection state corresponding to the user data to be selected; According to the selection state, the sample data is determined from the user data.
2. The user data extraction method according to claim 1, characterized in that: Determining the number of data to be selected based on the total number of data and the number of samples includes: When the total amount of data is greater than the sample amount, the non-sample amount is obtained according to the difference between the total amount of data and the sample amount; When the sample quantity is less than the non-sample quantity, determining the sample quantity as the to-be-selected quantity; When the sample quantity is not less than the non-sample quantity, the non-sample quantity is determined as the to-be-selected quantity.
3. The user data extraction method according to claim 1, characterized in that: The step size is determined according to the index of the user data, the offset, the total amount of data, and the amount to be selected, and the minimum index and the maximum index are determined according to the step size, the offset, the total amount of data, and the amount to be selected, including: Get the offset; When the index of the user data does not exceed the offset, setting the selection state corresponding to the user data to unselected; When the index of the user data exceeds the offset, determining a fixed tolerance according to the total amount of data and the amount to be selected; Determine a step size according to the index of the user data, the offset and the fixed tolerance; The minimum index and the maximum index are determined according to the step size, the fixed tolerance, and the offset.
4. The user data extraction method according to claim 2, characterized in that: Determining the sample data from the user data according to the selection state includes: When the sample quantity is less than the non-sample quantity, determining the user data in the selected state as the sample data; When the sample quantity is not less than the non-sample quantity, the user data whose selection status is unselected is determined as the sample data.
5. The user data extraction method according to claim 1, characterized in that: The obtaining corresponding user data according to the user information includes: Concurrently obtaining the user information from an extraction queue; the extraction queue is used to store user information corresponding to users participating in sample extraction; Determine a corresponding user queue according to the user information; the user information corresponds to the user queue one by one; the user queue is used to store user data corresponding to the user information; The user queue is traversed to obtain all user data corresponding to the user information.
6. The user data extraction method according to claim 1, characterized in that: Before the step of acquiring corresponding user data according to the user information, the method further includes: Receive a sample extraction request from a user; the sample extraction request includes a user identifier and a number of extractions; Generate the user information according to the user identifier, and add the user information to the extraction queue; Generate user data according to the user identifier and the extraction times, and add the user data to a user queue.
7. The user data extraction method according to claim 6, characterized in that: The generating user data according to the user identifier and the extraction times includes: When the sum of the current total number of extractions corresponding to the user identifier and the number of extractions does not exceed the extraction threshold, obtaining the minimum index value and the maximum index value corresponding to the sample extraction request according to the number of extractions; User data corresponding to the sample extraction request is generated according to the user identifier, the minimum index value and the maximum index value.
8. A user data extraction device, characterized in that: The device comprises: An acquisition module, used to acquire corresponding user data according to user information; the user information is used to characterize the users participating in sample extraction; A processing module, configured to determine the number to be selected according to the total number of data and the number of samples; the total number of data is the total number of all user data; the number of samples is the total number of sample data; a step length is determined according to the index of the user data, an offset, the total number of data and the number to be selected, and a minimum index and a maximum index are determined according to the step length, the offset, the total number of data and the number to be selected; the offset represents the starting index of the user data to be selected; when the index of the user data is the minimum index or the maximum index, the selection state corresponding to the user data is set to be selected; An extraction module is used to determine the sample data from the user data according to the selection state.
9. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the user data extraction method described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the user data extraction method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Lottery drawing data processing method and device, server and computer storage medium
CN107220853A
Data extracting method and device, electronic equipment and computer readable storage medium
CN108399266A
Lottery information determination method, device, apparatus, and storage medium
CN109542395A
Sample extraction method and device, equipment and computer storage medium
CN112686712A
Data processing method and device, electronic equipment and storage medium
CN113157743A