Data processing method and system

By simulating population evolution to generate target populations and automatically mining population behavior patterns, this technology solves the problems of low efficiency and insufficient accuracy of manual analysis in existing technologies, and achieves efficient and accurate population segmentation.

CN119719188BActive Publication Date: 2026-02-24ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411795837.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2026-02-24
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Existing technologies are inefficient in segmenting people, and manual summarization methods cannot accurately uncover complex behavioral patterns, resulting in poor service quality.

Method used

An initial behavioral sequence is generated by simulating population evolution. The target population is obtained through multiple generations of reproduction. Based on the target population, identification features are determined, and preset behavioral patterns are automatically mined.

Benefits of technology

It improves the accuracy and efficiency of population segmentation, reduces the subjectivity of human judgment, discovers multi-dimensional features and complex relationships, and enhances the representativeness of the identified features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719188B_ABST
    Figure CN119719188B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a data processing method and system. In the method, after obtaining historical behavior data corresponding to a plurality of users known to have a preset behavior pattern, M initial behavior sequences are generated based on the time sequence of each behavior in a plurality of behaviors of a user in the historical behavior data. By simulating the process of population evolution on the M initial behavior sequences, N target behavior sequences are obtained through multi-generation reproduction of the initial population, wherein the representation ability of the N target behavior sequences to the preset behavior model is higher than that of the M initial behavior sequences. In the above manner, automatic determination of mining the initial behavior sequences is realized to determine the target behavior sequences with higher representation ability to the preset behavior model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a data processing method and system. Background Technology

[0002] With the development of artificial intelligence technology, more and more fields have a need to provide personalized services to different groups of people. By analyzing user behavior data, users can be divided into different groups according to the needs of actual scenarios, thereby enabling the provision of different services to different groups.

[0003] Typically, to segment a user group, it's necessary to predefine the identification features corresponding to different groups. For example, for a group known to have a certain behavioral pattern, the behavioral data of users within that group is manually analyzed and patterns are summarized to obtain commonly occurring behavioral sequences. These behavioral sequences are then used as the identification features corresponding to that behavioral pattern. These features are then used to identify individuals exhibiting that behavioral pattern from a massive user base.

[0004] However, the above methods require manual analysis and summarization of large amounts of user behavior data, which is inefficient, especially when dealing with large datasets where processing speed is too slow. Furthermore, manual summarization typically only yields some simple, explicit characteristics of the population, and these characteristics have limited expressive power for behavioral patterns, leading to inaccurate population segmentation and thus impacting service quality.

[0005] The information in the background section is merely information known only to the inventor and does not imply that such information had entered the public domain before the filing date of this specification, nor does it imply that it can be considered prior art in this specification. Summary of the Invention

[0006] This specification provides a data processing method and system that can automatically mine the identification features corresponding to a preset behavioral pattern based on the historical behavioral data of a target user group with a known preset behavioral pattern by simulating population evolution.

[0007] Firstly, this specification provides a data processing method, comprising: obtaining historical behavior data corresponding to a target user group, the target user group including multiple users known to have preset behavior patterns; generating M initial behavior sequences based on the historical behavior data, wherein each initial behavior sequence represents multiple behaviors performed by a user in the target user group in chronological order, and M is an integer greater than 1; using the M initial behavior sequences as an initial population, and performing multi-generation breeding on the initial population by simulating a population evolution process to obtain a target population, the target population including N target behavior sequences, the N target behavior sequences having a higher representational ability of the preset behavior pattern than the M initial behavior sequences; and determining the identification features corresponding to the preset behavior pattern based on each target behavior sequence in the target population.

[0008] Secondly, this specification also provides a data processing system, including at least one storage medium and at least one processor, wherein the at least one storage medium stores at least one instruction set for performing a data processing method; the at least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set during operation and executes the data processing method described in any of the first aspects above according to the instructions of the at least one instruction set.

[0009] Thirdly, this specification also provides a computer-readable non-volatile storage medium, wherein the computer-readable non-volatile storage medium stores at least one instruction set, which, when executed by at least one processor, implements the data processing method as described in any of the first aspects above.

[0010] As can be seen from the above technical solutions, the data processing method and system provided in this specification, after acquiring historical behavior data of a target user group with a known preset behavior pattern, generates M initial behavior sequences based on the historical behavior data. These M initial behavior sequences are used as an initial population, and the initial population is multi-generationally reproduced through a simulated population evolution process to obtain the target population. The behavior sequences in the target population have a higher representational ability of the preset behavior pattern than the initial behavior sequences, ensuring that the target population can accurately represent the preset behavior pattern. Furthermore, the multi-generational reproduction method can generate diverse behavior sequences, ensuring that the behavioral characteristics of different users can be captured. Population evolution can discover implicit features and complex relationships between features across multiple dimensions. Based on each target behavior sequence in the target population, the identification features of the preset behavior pattern are determined. Identifying these features allows for more accurate matching of user groups with preset behavior patterns, improving the accuracy of subsequent user group segmentation. Moreover, all of the above data processing methods are automated, improving data processing efficiency and reducing the subjectivity and bias of human judgment.

[0011] The data processing methods and other functions of the system provided in this specification are partially listed in the following description. The inventive aspects of the data processing methods and systems provided in this specification can be fully explained through practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A schematic diagram of an application scenario provided according to an embodiment of this specification is shown;

[0014] Figure 2 A schematic diagram of the hardware structure of a computing device provided according to some embodiments of this specification is shown;

[0015] Figure 3 A schematic flowchart of a data processing method according to an embodiment of this specification is shown;

[0016] Figure 4 A schematic flowchart of a reproduction process according to another embodiment of this specification is shown;

[0017] Figure 5 A schematic diagram of a multigenerational reproduction process provided according to an embodiment of this specification is shown;

[0018] Figure 6 A schematic diagram of a single-point crossover algorithm provided according to an embodiment of this specification is shown;

[0019] Figure 7 A schematic diagram of a multi-point crossover algorithm provided according to an embodiment of this specification is shown; and

[0020] Figure 8 A schematic diagram of a mutation algorithm provided according to an embodiment of this specification is shown. Detailed Implementation

[0021] The following description provides specific application scenarios and requirements for this specification, intended to enable those skilled in the art to make and use the contents of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.

[0022] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. When used in this specification, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.

[0023] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0024] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.

[0025] In this specification, "X includes at least one of A, B, or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is, X may include only one of A, B, and C, or any combination of A, B, and C, as well as other possible content / elements. The arbitrary combination of A, B, and C can be A, B, C, AB, AC, BC, or ABC.

[0026] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0027] It should be noted that the user data obtained in this manual is authorized by the user and does not involve user privacy.

[0028] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user-related information involved in the technical solutions of this specification all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0029] For ease of description, the terms that will appear later in this manual will be explained first.

[0030] Genetic Algorithm (GA) is a search and optimization algorithm inspired by natural selection and genetic mechanisms. It searches for the optimal or near-optimal solution to a problem by simulating operations such as selection, crossover (recombination), and mutation in biological evolution.

[0031] User behavior refers to the various activities and interaction patterns exhibited by users in a specific environment or system. These behaviors can include various operations and interactions of users in digital environments such as websites, applications, social media platforms, and e-commerce platforms.

[0032] A behavior sequence refers to a series of actions performed by a user in a specific environment or system in chronological order. These actions can include clicking, browsing, purchasing, searching, sharing, and other operations. By recording the temporal order of these actions, an ordered behavior sequence can be formed.

[0033] Event tracking: This is a data collection technique that involves inserting code (i.e., "event points") at key locations in an application or website to record user behavior data at specific points. This data can include user actions such as clicks, swipes, input, and page dwell time, as well as basic user information (such as user ID, device information, etc.).

[0034] Crossover: In genetic algorithms, crossover is an important operation used to generate new offspring. Crossover combines the genes of two parent individuals to produce new offspring.

[0035] The technical solutions provided in this manual can be applied to scenarios requiring audience segmentation. For example, in personalized recommendation scenarios, segmenting audiences allows for the recommendation of different content based on the preferences of different groups. As another example, in risk control scenarios, audience segmentation allows for the precise identification of high-risk individuals, enabling targeted prevention and control measures for these groups.

[0036] Taking personalized recommendation scenarios as an example, to segment users, it's necessary to predefine the identification features corresponding to different user groups. For instance, for a group known to have a certain behavioral pattern, the behavioral data of users within that group is manually analyzed and patterns are summarized to obtain the common behavioral sequences among users in that group. These behavioral sequences are then used as the identification features corresponding to that behavioral pattern. These features are then used to identify users with that behavioral pattern from a massive user base. This method of manually analyzing and summarizing behavioral sequences is not only costly in terms of manpower and time, but also slow. It fails to achieve truly personalized recommendations and ignores the logical relationship between user behavior and outcomes. Furthermore, manual processing can only discover strongly correlated sequences, making it difficult to uncover weakly correlated but effective behavioral sequences. This lack of depth in understanding user needs may lead to inaccurate analysis results, affecting the quality of subsequent decision-making.

[0037] The technical solution provided in this specification, after acquiring historical behavior data corresponding to the target user group, can generate M initial behavior sequences based on the historical behavior data, and use these M initial behavior sequences as the initial population. Each behavior sequence represents multiple behaviors performed by a user in chronological order. This method of determining the initial population based on historical behavior data ensures the diversity of the initial population. Subsequently, the technical solution provided in this specification simulates the population evolution process to propagate the initial population through multiple generations, retaining the better-performing target behavior sequences and removing the poorly performing ones, resulting in a target population with a higher representational ability of the preset behavior pattern than the initial population. Finally, the technical solution provided in this specification also determines the identification features corresponding to the preset behavior pattern based on each target behavior sequence in the target population. This method of determining identification features based on the target population ensures that the obtained identification features are highly representative of the preset behavior pattern because the target population has a strong representational ability of the preset behavior. In other words, the technical solution provided in this specification not only improves the efficiency and accuracy of identification feature determination, but also ensures that the obtained identification features better reflect the user's behavior pattern.

[0038] It should be noted that the above description of application scenarios is only one of the many usage scenarios provided in this specification. Those skilled in the art should understand that when the data processing methods and systems provided in this specification are applied to other usage scenarios, their implementation methods and technical effects are similar.

[0039] Figure 1 A schematic diagram of an application scenario 100 provided according to an embodiment of this specification is shown.

[0040] like Figure 1As shown, application scenario 100 may include a feature mining system 130 and a user mining system 150. Application scenario 100 involves two execution phases. In the first phase, the feature mining system 130 analyzes the historical behavioral data of multiple users (target user groups) known to have preset behavioral patterns to mine the identification features corresponding to these preset behavioral patterns. Preset behavioral patterns refer to recurring behavioral patterns or sequences of behavior in a certain context. For example, preset behavioral patterns may include, but are not limited to: buying a cup of coffee after breakfast, exercising during lunchtime, habitually buying high-value items, habitually buying discounted items, and frequently making large-amount transfers. In the second phase, the user mining system 150 mines user groups with the preset behavioral patterns from the user set based on the aforementioned identification features, thereby achieving the selection (mining) of users with preset behavioral patterns based on identification features. It is understood that the target user group may consist of a small number of users known to have preset behavioral patterns, while the user group mined by the user mining system 150 may consist of a larger number of users with preset behavioral patterns.

[0041] The feature mining system 130 can be a computing system with a certain computing capability. The feature mining system 130 can execute the data processing methods provided in this specification, thereby extracting identification features corresponding to preset behavioral patterns from historical behavioral data corresponding to a target user group. The feature mining system 130 can store data or instructions for executing the data processing methods described in this specification, and can execute or be used to execute the data or instructions. The feature mining system 130 may include hardware devices with data information processing functions and the necessary programs required to drive the hardware devices.

[0042] The feature mining system 130 can be a single computing device or a cluster system composed of multiple computing devices. The data or instructions stored in the feature mining system 130 for executing the data processing methods described in this specification can adopt any form of system architecture. For example, layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc.

[0043] In some embodiments, the feature mining system 130 can first generate M initial behavior sequences based on historical behavior data corresponding to the target user group, and use these M initial behavior sequences as an initial population. Subsequently, the feature mining system 130 performs multi-generation breeding on the initial population by simulating a population evolution process to obtain a target population. The target population includes N target behavior sequences, and the representational ability of the N target behavior sequences for the preset behavior pattern is higher than that of the M initial behavior sequences. Finally, the feature mining system 130 can determine the recognition features corresponding to the preset behavior pattern based on each target behavior sequence in the target population.

[0044] In other words, the feature mining system 130 can mine the identification features corresponding to preset behavioral patterns based on the historical behavioral data of the target user group. Compared to the manual method of sorting and segmenting users based on their historical behavioral data, the technical solution provided in this specification does not require manual sorting for user mining. The user mining system 150 can directly mine users with preset behavioral patterns from the user set. Subsequently, the user mining system 150 automatically mines users with preset behavioral patterns from the user set based on the identification features. Therefore, the technical solution provided in this specification improves the speed of determining the target population, thereby improving the speed of determining the identification features corresponding to users with preset behavioral patterns and reducing the labor and time costs caused by manual sorting. Furthermore, in the embodiments of this specification, the user mining system 150 determines the target behavioral sequence through multi-generation breeding, which can not only discover strongly correlated behavioral sequences, but also some weakly correlated behavioral sequences that have a good representation ability of preset behavioral capabilities. This improves the accuracy and diversity of target behavioral sequences contained in the target population, thereby improving the accuracy of the identification features corresponding to the preset behavioral patterns determined based on the target population.

[0045] Furthermore, since the features mined by the feature mining system 130 are the corresponding features under the preset behavioral pattern, they are highly representative of the preset behavioral pattern and have specificity. This avoids the influence of other features (features unrelated to the preset behavioral pattern) on user mining, thus making the users with the preset behavioral pattern mined from the user set more accurate.

[0046] The user mining system 150 and the feature mining system 130 may correspond to the same computing system or to different computing systems; this specification does not limit this.

[0047] In some embodiments, the data processing system provided in this specification may correspond to Figure 1 The feature mining system 130 in the specification. In some embodiments, the data processing system provided in this specification may include... Figure 1 The feature mining system 130 and user mining system 150 are included.

[0048] Figure 2 A schematic diagram of the hardware structure of a computing device 200 according to some embodiments of this specification is shown. This computing device 200 can be used as... Figure 1 The feature mining system 130 is described in some embodiments. When the feature mining system 130 employs a device cluster, the computing device 200 can be any one of the devices in the feature mining system 130.

[0049] like Figure 2 As shown, the computing device 200 includes at least one storage medium 230 and at least one processor 220. In some embodiments, the computing device 200 may further include an internal communication bus 210. In some embodiments, the computing device 200 may further include a communication port 250. In some embodiments, the computing device 200 may further include I / O components 260.

[0050] The internal communication bus 210 can connect different system components, including storage medium 230 and processor 220. I / O component 260 supports input / output between computing device 200 and other components.

[0051] Communication port 250 is used for data communication between computing device 200 and the outside world. For example, computing device 200 can connect to a network through communication port 250.

[0052] Storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 236. Storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set is computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc., that execute the data processing methods provided in this specification.

[0053] At least one processor 220 is communicatively connected to at least one storage medium 230 via an internal communication bus 210. The at least one processor 220 is used to execute at least one instruction set. When the system 130 is running, the at least one processor 220 reads at least one instruction set and executes the data processing methods provided in this specification according to the instructions of the at least one instruction set.

[0054] Processor 220 can execute all the steps included in the data processing method. Processor 220 can be in the form of one or more processors. Processor 220 can issue execution instructions. Processor 220 may include one or more hardware processors, such as microcontrollers, microprocessors, reduced instruction set computers (RISC), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), central processing units (CPUs), graphics processing units (GPUs), physical processing units (PPUs), microcontroller units, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), advanced RISC machines (ARMs), programmable logic devices (PLDs), any circuit or processor capable of performing one or more functions, or any combination thereof.

[0055] For illustrative purposes only, only one processor 220 is shown in the accompanying drawings of the computing device 200. However, it should be noted that the computing device 200 may also include multiple processors. Therefore, the operation and / or method steps disclosed herein may be executed by a single processor or by multiple processors in combination, as described herein. For example, if processor 220 of the computing device 200 in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 220 (e.g., a first processor executes step A, a second processor executes step B, or the first and second processors jointly execute steps A and B).

[0056] Figure 3 A schematic flowchart of a data processing method P300 according to an embodiment of this specification is shown. This data processing method P300 can be executed by a feature mining system 130. Figure 3 As shown, the method P300 provided in this specification may include S310-S370, wherein:

[0057] S310: Obtain historical behavioral data corresponding to the target user group, which includes multiple users known to have preset behavioral patterns.

[0058] In the embodiments of this specification, preset behavioral patterns may include: frequently engaging in high-risk behaviors, habitually purchasing discounted goods, habitually purchasing high-value goods, and habitually drinking coffee every morning. It should be understood that the above embodiments are merely illustrative, and the specific content of the preset behavioral patterns may be flexibly adjusted according to user needs, and is not limited to what is given in the above embodiments. Application scenarios may also include health management, online education, or content consumption, etc., and will be determined based on the actual application situation.

[0059] When determining the historical behavioral data corresponding to a target user group, the relevant historical behavioral data can be determined based on the current application scenario and the target user group. For example, historical behavioral data in an e-commerce scenario can be user behavior logs. Historical behavioral data in a social media scenario can be user interaction records. Historical behavioral data in a financial platform scenario can be user transaction records, etc. Obtaining historical behavioral data of a target user group includes, but is not limited to, obtaining user identifier (ID) information, behavior type, behavior time, behavior object, etc.

[0060] S330: Generate M initial behavior sequences based on historical behavior data, where each initial behavior sequence represents multiple behaviors performed by a user in the target user group in chronological order, and M is an integer greater than 1.

[0061] In practical applications, user historical behavior data is typically collected from pre-deployed data collection points within the code. These data collection points are called event tracking points. For example, in a social media scenario, an event tracking point can be pre-set in the code segment of the social application used to respond to a "favorite" action. When this code segment is triggered, it indicates that the user has performed a "favorite" action, and this action is recorded as user behavior data. The data collected by each event tracking point can be applied to one user action. Arranging the data collected from different event tracking points in chronological order yields the user's behavior sequence.

[0062] In some possible implementations, a large number of data collection points (tracking points) may be widely deployed in an application scenario (application or website) to record various user behaviors. These tracking points can cover almost all user interactions during use. However, some tracking points among the massive amount of data may not be representative enough of the user's key behavioral capabilities and may not be highly correlated with the user's behavioral patterns. Therefore, to reduce the impact of irrelevant data, some representative and valuable key tracking points can be pre-determined from the massive amount of data based on preset behavioral patterns.

[0063] For example, in a shopping scenario, a large number of event tracking points may include: page loading tracking points, login / registration tracking points, personal profile modification tracking points, shipping address filling tracking points, page scrolling tracking points, search submission tracking points, product description reading tracking points, product click tracking points, adding to cart click tracking points, buying now click tracking points, navigation menu click tracking points, promotional activity click tracking points, product sharing click tracking points, viewing similar products click tracking points, etc.

[0064] If the target user group has a pre-defined behavior pattern of purchasing high-value goods on e-commerce platforms, then the page loading tracking (which is highly correlated with the user's current network conditions), login / registration tracking, shipping address entry tracking, and personal profile modification tracking are not highly correlated with the pre-defined behavior pattern, and therefore it is not necessary to obtain the behavioral data corresponding to the above tracking points.

[0065] In other words, in the above scenario, the key event tracking points (number L) selected from the massive event tracking points may include: page scrolling event tracking points, search submission event tracking points, product description reading event tracking points, add to cart click event tracking points, buy now click event tracking points, promotional activity click event tracking points, product sharing click event tracking points, and viewing similar products click event tracking points, etc. It should be understood that the types of event tracking points and the selection of key event tracking points may be included in the above massive event tracking points can be flexibly adjusted according to the current application scenario, user needs, and actual conditions, and are not limited to those given in the above embodiments.

[0066] To ensure the validity of the data, the historical behavior data can be preprocessed before extracting the initial behavior sequence. This process removes invalid, redundant, and abnormal data from the historical behavior data, ensuring that the extracted data is accurate and complete.

[0067] In some possible implementations, a user may generate a large amount of behavioral data when using an application or website, which may include many irrelevant or repetitive behaviors. Analyzing this data based on the large amount of behavioral data from each user in the target user group is not only computationally inefficient and computationally intensive, but also carries the risk of overfitting. Therefore, to improve computational efficiency, behavioral data from M users can be randomly selected from the historical behavioral data of a large target user group. From this data, a set of key behaviors corresponding to a preset behavioral pattern can be obtained (the set of key behaviors is determined based on key tracking points; one tracking point corresponds to one behavior, and multiple key tracking points correspond to multiple key behaviors, constituting the set of key behaviors). The set of key behaviors includes multiple key behaviors related to the preset behavioral pattern. Subsequently, for each of the M users, multiple behaviors of that user that fall within the set of key behaviors are selected from the user's behavioral data, and an initial behavior sequence is generated based on these multiple behaviors.

[0068] In other words, the embodiments in this specification first randomly select behavioral data from M users from a large amount of historical behavioral data of the target user group. Compared to directly obtaining behavioral data from all target user groups, this reduces the workload in subsequent data processing to some extent. Furthermore, the technical solution provided in this specification does not obtain complete behavioral data for each user, but rather filters the behavioral data for each user. Given a large amount of behavioral data for a single user, it only obtains a set of key behaviors corresponding to preset behavioral patterns based on key tracking points. This further reduces the workload in subsequent data processing to some extent.

[0069] To further improve the efficiency of subsequent data processing, in one embodiment of this specification, behavioral data from all key behavior sets is not acquired. Instead, when determining an initial behavior sequence, multiple key behaviors are extracted from the key behavior set for each of the M users. These behaviors are then sorted based on their chronological order to form the initial behavior sequence for that user. In other words, an initial behavior sequence may include P behaviors executed by a user in chronological order, where each of the P behaviors comes from the key behavior set, and P is an integer greater than 1 and less than L (the number of key tracking points). This method of extracting multiple behaviors from the key behavior set not only reduces data noise but also reduces the workload of subsequent data processing, improving computational efficiency.

[0070] In the embodiments of this specification, to ensure the diversity of the initial population and to guarantee that the behavior sequence can contain more different types of behaviors, the method of selecting P key tracking points can be random sampling. That is, based on the historical behavior data of each user in the target user group, P key behaviors are randomly selected from the key behavior set corresponding to each user and the preset behavior pattern. The key behaviors are then sorted according to their execution time to obtain the initial behavior sequence for each user. This method of extracting multiple key behaviors ensures the representativeness of the obtained behavior sequence because all key behaviors are extracted from the key behavior set. The random sampling of multiple key behaviors ensures the diversity of the obtained behavior sequence, thereby ensuring the quality of the obtained initial population.

[0071] The following example illustrates key event tracking (number L): page scrolling tracking, search submission tracking, product description reading tracking, adding to cart clicking tracking, "buy now" clicking tracking, promotional activity clicking tracking, product sharing clicking tracking, and viewing similar products clicking tracking. For instance, P might be 3, meaning each randomly selected P tracking points could be any 3 of the L key event tracking points mentioned above. For example, M initial behavior sequences might include Initial Behavior Sequence 1, Initial Behavior Sequence 2, and Initial Behavior Sequence 3. Initial Behavior Sequence 1 consists of the actions of searching for a product, adding to cart, and clicking "buy now." Initial Behavior Sequence 2 consists of viewing product details, sharing a product, and clicking "buy now." Initial Behavior Sequence 3 consists of adding to cart, clicking a product, and sharing a product. It should be understood that the above embodiments are merely illustrative. The specific settings for the number of L and P, as well as the actual situation of the tracking points, can be flexibly adjusted according to user needs and are not limited to the embodiments described above.

[0072] In the embodiments described in this specification, after generating M initial behavior sequences based on historical behavior data, each initial behavior sequence can be encoded with real numbers to facilitate subsequent calculations and processing, and to facilitate subsequent operations such as crossover and mutation. Real number encoding can compress multidimensional discrete data into one-dimensional or low-dimensional continuous data, reducing the problem of dimensionality explosion, reducing the complexity of subsequent calculations, and improving computational efficiency.

[0073] In the embodiments of this specification, a corresponding real-number code is pre-assigned for each behavior to ensure that the representation of the real-number code for each behavior is unique in the behavior sequence, facilitating subsequent calculation and analysis. One way to assign a corresponding real-number code to each behavior is to pre-create a mapping table. For example, in a shopping scenario, each behavior and its corresponding real-number code are shown in Table 1 below.

[0074] Table 1

[0075] Behavior Real number encoding Search for products 1 Click on the product 2 View product details 3 Share products 4 add to the cart 5 Visit shopping cart page 6 Click "Buy Now" 7 Fill in the shipping address 8 Confirm payment 9 Payment successful 10

[0076] The encoding method involves converting each action data in the initial action sequence into a real number according to a preset mapping relationship, thus transforming each initial action sequence into a real number sequence. Taking initial action sequence 1, initial action sequence 2, and initial action sequence 3 as examples, the real number sequence corresponding to initial action sequence 1 [search for product, add to cart, click "Buy Now"] is [1, 5, 7]. The real number sequence corresponding to initial action sequence 2 [view product details, share product, click "Buy Now"] is [3, 4, 7]. The real number sequence corresponding to initial action sequence 3 [add to cart, click product, share product] is [5, 2, 4]. By converting the string sequence into a real number sequence in this way, the data format in the action sequence is unified, the efficiency of data processing is improved, and the complexity and inefficiency of string operations are avoided. It should be understood that the above embodiments are merely illustrative and are not limited to the embodiments given above.

[0077] S350: Using M initial behavior sequences as the initial population, the target population is obtained by simulating the population evolution process and breeding the initial population for multiple generations. The target population includes N target behavior sequences. The N target behavior sequences have a higher representation ability of the preset behavior pattern than the M initial behavior sequences.

[0078] In the embodiments of this specification, in order to ensure the accuracy of the target population, it can be continuously iterated through multiple generations of breeding until the target population is obtained.

[0079] Figure 4 A schematic flowchart of a reproduction process according to another embodiment of this specification is shown, such as... Figure 4 As shown, multi-generational reproduction may consist of a total of X generations, with each generation corresponding to one generation of the population. During the reproduction process, the ability of the behavioral sequences within the population to represent the preset behavioral pattern gradually increases until the target population is obtained. The target population's ability to represent the preset behavioral model is higher than that of the initial population.

[0080] Figure 5 A schematic diagram of a multigenerational reproduction process according to an embodiment of this specification is shown, such as... Figure 5 As shown, the i-th generation of reproduction in multi-generational breeding includes S351-S355, where:

[0081] S351: Select a subset of behavioral sequences that meet the evolutionary conditions from the current population to form the first set of behavioral sequences. The evolutionary conditions are related to the ability to represent the preset behavioral patterns.

[0082] The fitness of each behavior sequence can be calculated based on a fitness function. The fitness function uses the behavior sequence as the independent variable and fitness as the dependent variable. The fitness obtained from the fitness function characterizes the ability of a behavior sequence to represent a preset behavior pattern. Different preset behavior patterns correspond to different fitness functions. For example, in a lending scenario, the optimization problem is to increase the borrowing rate, so the fitness function should be positively correlated with the borrowing rate. If the current application scenario is a shopping scenario, the optimization problem is to increase the purchase rate, so the fitness should be positively correlated with the purchase rate. That is, in the embodiments of this specification, the fitness function corresponding to the preset behavior pattern is first obtained, and then the fitness corresponding to each behavior sequence in the current population is determined based on the fitness function. Based on the fitness corresponding to each behavior sequence in the current population, a subset of behavior sequences that meet the evolutionary conditions are selected from the current population to form a first set of behavior sequences.

[0083] In some possible embodiments, the evolutionary condition can be a preset fitness threshold. Then, the way to select a portion of the behavioral sequences from the current population based on fitness can be to select a portion of the behavioral sequences from the current population whose fitness meets the preset fitness threshold.

[0084] To avoid excessively large differences in fitness values ​​between behavioral sequences, which could lead to some highly fit behavioral sequences rapidly dominating the population, resulting in decreased genetic diversity and premature convergence to a local optimum, the embodiments in this specification use an exponential transformation to scale the fitness function, standardizing the fitness to a uniform range and reducing the problem of excessively large differences in fitness between different behavioral sequences.

[0085] After setting the fitness function, each behavior sequence in the current population is evaluated using the fitness function to quantify the quality of user behavior in each set of behavior sequences. Higher fitness indicates a better representation of the preset behavior pattern by the set of behavior sequences. Lower fitness indicates a worse representation of the preset behavior pattern by the set of behavior sequences.

[0086] In the process of multigenerational reproduction, some infeasible solutions may exist. That is, some behavioral sequences, although satisfying the preset fitness threshold, violate the constraints of reality. Therefore, such behavioral sequences cannot be considered to satisfy the preset fitness threshold. For example, behavioral sequence 7 [sharing a product, filling in a shipping address, searching for a product] may have a high fitness value and satisfy the fitness threshold. However, when the preset behavioral pattern is to purchase high-value products, behavioral sequence 7 obviously violates reality. Few people would purchase high-value products after performing the above behavior, so it has no practical significance.

[0087] To address the aforementioned issues, this specification considers the feasibility of behavioral sequences in its embodiments. By introducing a penalty term, configured to adjust the fitness of a behavioral sequence based on its feasibility, the fitness of infeasible solutions (e.g., behavioral sequence 7 mentioned above) is significantly lower than that of any feasible solutions. When the feasibility of the first behavioral sequence is lower than that of the second behavioral sequence, the adjusted fitness corresponding to the first behavioral sequence is lower than that corresponding to the second behavioral sequence. For example, the fitness value of the penalty term can be set to a very low preset fitness value. This method of introducing a penalty term significantly reduces the fitness value of infeasible solutions during subsequent generations of reproduction, prompting them to be gradually eliminated or transformed into feasible solutions.

[0088] In some possible embodiments, the evolutionary condition can also be a preset probability threshold. The evolution of each behavioral sequence in the current population can be determined using a Roulette Wheel Selection method. The core idea of ​​this method is to determine the probability of each behavioral sequence being selected for the next generation based on its fitness. Behavioral sequences with higher fitness have a greater probability of being selected, and vice versa. Specifically, assuming the current population contains K behavioral sequences, the i-th behavioral sequence a... i The probability p of being selected i It is the sequence of actions a i fitness f i The sum of fitness of all individuals in the current population The ratio. In other words,

[0089] When selecting behavioral sequences that satisfy evolutionary conditions, one approach is to randomly select from multiple behavioral sequences that satisfy the conditions to determine the subset of behavioral sequences that meet the evolutionary conditions. Alternatively, multiple behavioral sequences can be sorted according to their selection probabilities, and the top-ranked behavioral sequences can be determined based on the sorting results. The number of subset behavioral sequences can be, for example, a preset constant value, or a preset proportion, the specific number determined based on the number of behavioral sequences in the current population and the preset proportion. It should be understood that the above embodiments are merely illustrative, and the specific method for determining the subset of behavioral sequences that satisfy evolutionary conditions can be flexibly adjusted according to user needs and is not limited to the methods given in the above embodiments.

[0090] In the embodiments of this specification, the mean fitness of the current population can be determined based on the fitness of each behavioral sequence in the current population. Based on the fitness mean and each behavioral sequence a in the current population i fitness f i The ratio determines the selection probability p of each action sequence.i In other words, Behavioral sequences with a selection probability greater than a preset probability threshold are selected from the current population to form a first set of behavioral sequences. It should be understood that the above embodiments are merely illustrative, and the specific evolutionary conditions and methods for determining the selection probability of each behavioral sequence can be flexibly adjusted according to user needs and are not limited to those given in the above embodiments.

[0091] S353: Perform crossover operations on different behavior sequences in the first behavior sequence set to obtain multiple evolutionary behavior sequences, and add the multiple evolutionary behavior sequences to the first behavior sequence set to obtain the second behavior sequence set.

[0092] Since the behavioral sequences contained in the first behavioral sequence set all have high fitness and meet the evolutionary conditions, it indicates that the behavioral sequences in the first behavioral sequence set better represent the preset behavioral pattern. These superior behavioral sequences should be inherited by the next generation. Therefore, in the embodiments provided in this specification, the first behavioral sequence set is directly copied into the second behavioral sequence set, and a crossover operation is performed based on the first behavioral sequence set to obtain multiple evolutionary behavioral sequences. In other words, the second behavioral sequence set is composed of the first behavioral sequence set and the multiple evolutionary behavioral sequences obtained after crossover based on the first behavioral sequence set.

[0093] In the embodiments of this specification, the evolutionary behavior sequence includes multiple sub-sequences, wherein each sub-sequence originates from at least two different behavior sequences in the first behavior sequence set. That is, there are at least two behavior sequences participating in a crossover operation, and there can be three or even more. The specific number of behavior sequences participating in a crossover operation can be flexibly adjusted according to user needs and is not limited to the examples given above.

[0094] In the following embodiments, the example of two behavior sequences participating in a crossover operation is used for illustration: For example, two behavior sequences can be randomly selected from the first behavior sequence as parent behavior sequences, and some behavior data in the parent behavior sequences can be exchanged according to a preset crossover rate to obtain child behavior sequences, and the obtained child behavior sequences can be added to the second behavior sequence set.

[0095] Based on the number of intersection points, crossover operations can be divided into single-point crossover operations and multi-point crossover operations. Figure 6 A schematic diagram of a single-point crossover algorithm according to an embodiment of this specification is shown. Figure 6 As shown, the single-point crossover operation can randomly determine a crossover point in the parent behavior sequence according to the preset crossover rate, and exchange some behavior data in the two parent behavior sequences at the crossover point to generate two child behavior sequences.

[0096] like Figure 6 As shown, an exemplary description of a single-point crossover operation can be as follows: Assume there are two parent behavior sequences A and B, whose behavior data are: Parent behavior sequence A (behavior 1, behavior 2, behavior 3, behavior 4, behavior 5, behavior 6, behavior 7, behavior 8); Parent behavior sequence B (behavior 8, behavior 7, behavior 6, behavior 5, behavior 4, behavior 3, behavior 2, behavior 1). The position after the second behavior in the sequence is taken as the crossover point, and the behavior data before (or after) the crossover point is swapped. This generates two child behavior sequences: Child behavior sequence 1 (behavior 8, behavior 7, behavior 3, behavior 4, behavior 5, behavior 6, behavior 7, behavior 8); Child behavior sequence 2 (behavior 1, behavior 2, behavior 6, behavior 5, behavior 4, behavior 3, behavior 2, behavior 1).

[0097] Figure 7 A schematic diagram of a multi-point crossover algorithm according to an embodiment of this specification is shown. For example... Figure 7 As shown, the multi-point crossover operation can select multiple crossover points in the parent behavior sequence, and then exchange part of the behavior data of the two parent behavior sequences at the multiple crossover points to generate two new child behavior sequences.

[0098] Using the two parent behavior sequences A and B as examples, a multi-point crossover operation can be implemented, for instance, by taking the positions after the second and fifth behaviors in the sequence as crossover points. After swapping the behavior data outside these two crossover points (or between them), two child behavior sequences are generated: Child Behavior Sequence 3 (behavior 8, behavior 7, behavior 3, behavior 4, behavior 5, behavior 3, behavior 2, behavior 1) and Child Behavior Sequence 4 (behavior 1, behavior 2, behavior 6, behavior 5, behavior 4, behavior 6, behavior 7, behavior 8). It should be understood that the above embodiments are merely illustrative. In the embodiments described in this specification, only single-point crossover operations, only multi-point crossover operations, or a combination of single-point and multi-point crossover operations can be used. The specific crossover method can be flexibly adjusted according to the application scenario and user needs, and this specification does not impose any restrictions.

[0099] S355: Generate the population after the i-th generation of reproduction based on the set of sequences in the second row, where i is any natural number. When i = 1, the current population is the initial population, and when i > 1, the current population is the population after the (i-1)-th generation of reproduction.

[0100] In the embodiments described herein, the population generated after the i-th generation can be generated by performing mutation operations on a portion of the behavior sequences in the second behavior sequence set to obtain multiple mutated behavior sequences. These multiple mutated behavior sequences, along with other unmutated behavior sequences from the second behavior sequence set, are used as the population generated after the i-th generation.

[0101] The mutation operation can be one or more of the following: randomly adding a preset number of behaviors to the behavior sequence; randomly deleting a preset number of behaviors from the behavior sequence; or randomly replacing some behaviors in the behavior sequence with other behaviors. The preset number may be determined by the mutation rate; a higher mutation rate results in a larger preset number, and a lower mutation rate results in a smaller preset number. Mutating a portion of the behavior sequence can create new behavior sequences, preventing the algorithm from shrinking to a local optimum. The specific mutation rate setting can be flexibly adjusted according to user needs and is not limited to the examples given above.

[0102] For example, taking the parent behavior sequence C (behavior 1, behavior 2, behavior 6, behavior 5, behavior 3) as an example, for a mutation operation that randomly adds a preset number of behaviors, the mutation method may be: randomly selecting one position in the parent behavior sequence to add a new behavior. In one embodiment, a random behavior may be added after the first behavior in the parent behavior sequence C through a mutation operation, resulting in the mutated child behavior sequence C1 (behavior 1, behavior 5, behavior 2, behavior 6, behavior 5, behavior 3).

[0103] For a mutation operation that randomly deletes a preset number of behaviors from a behavior sequence, the mutation method may be as follows: randomly delete behaviors at two positions in the parent behavior sequence. In one embodiment, the mutation operation may randomly delete the first and third behaviors in the parent behavior sequence C to obtain the mutated child behavior sequence C1 (behavior 2, behavior 5, behavior 3).

[0104] Figure 8 A schematic diagram of a mutation algorithm provided according to an embodiment of this specification is shown. Figure 8 The mutation operation shown in the diagram randomly replaces some behaviors in a sequence of behaviors with other behaviors. For example... Figure 8As shown, taking the parent behavior sequence C as an example, its mutation method may be to randomly replace the behaviors in two parent behavior sequences with other random behaviors. In one embodiment, the first and fourth behaviors in the parent behavior sequence C may be randomly mutated, resulting in the mutated child behavior sequence C3 (behavior 9, behavior 2, behavior 6, behavior 8, behavior 3). It should be understood that the above embodiments are only illustrative examples, and the specific preset quantity settings, the positions of the randomly determined mutated behaviors, etc., can be adjusted according to the actual situation or user needs, and are not limited to the embodiments given above.

[0105] In the embodiments of this specification, the termination condition for multi-generation breeding includes at least one of the following: the number of iterations already bred is greater than or equal to a preset number of iterations T. max In other words, the current number of iterations T that has been propagated is greater than or equal to the preset number of iterations T. max (T>T max When this happens, multigenerational reproduction ends.

[0106] The termination condition for multi-generation reproduction can also be: the representational ability of the behavioral sequences contained in the current population to the preset behavioral pattern reaches a convergence condition. The convergence condition can be represented by introducing an evaluation function. When the change in the evaluation function is less than a preset threshold, the representational ability of the current population can be considered converged, and multi-generation reproduction ends. It should be understood that the above embodiments are merely illustrative examples. The specific termination condition for multi-generation reproduction can also be that each behavioral sequence in the current population reaches a preset fitness threshold, or that the diversity of the current population decreases to a preset threshold, etc. The specific termination condition for multi-generation reproduction can be flexibly adjusted according to user needs and is not limited to the conditions given in the above embodiments.

[0107] S370: Based on the target behavior sequences in the target population, determine the identification features corresponding to the preset behavior patterns.

[0108] After identifying the identifying features, users with behavioral patterns can be extracted from the user set based on these features. For example, for each user in the user set, behavioral characteristics can be determined based on that user's historical behavioral data, and the similarity between those behavioral characteristics and the identified features can be determined. If the similarity is greater than or equal to a preset threshold, the user is determined to have the preset behavioral pattern. If the similarity is less than the preset threshold, the user is determined not to have the preset behavioral pattern. Users with preset behavioral patterns can be extracted from the user set in this way.

[0109] In personalized recommendation applications, identifying more users with pre-defined behavioral patterns from the user set facilitates more precise recommendations, enabling personalized or targeted recommendations. In marketing applications, the above methods can achieve user segmentation, facilitating targeted marketing strategies. In other potential applications, different user management or maintenance strategies can be adopted based on different user behavioral patterns to improve user retention and repurchase rates. Alternatively, in risk identification scenarios, identifying the characteristics of abnormal behavioral patterns can help detect and prevent potential fraudulent activities in a timely manner. It should be understood that the subsequent processing of users with behavioral patterns identified from the user set can be flexibly adjusted according to the current application scenario and the actual needs of users, and is not limited to the embodiments given above.

[0110] In summary, the data processing method and system provided in this specification, after acquiring historical behavioral data of a target user group with a known preset behavioral pattern, generates M initial behavioral sequences based on the historical behavioral data. These M initial behavioral sequences are used as an initial population, and the initial population is multiplied through multiple generations by simulating a population evolution process to obtain the target population. The behavioral sequences in the target population have a higher representational ability of the preset behavioral pattern than the initial behavioral sequences, ensuring that the target population can accurately represent the preset behavioral pattern. Furthermore, the multi-generational reproduction method can generate diverse behavioral sequences, ensuring the capture of behavioral characteristics of different users. Population evolution can discover implicit features and complex relationships between features across multiple dimensions. Based on each target behavioral sequence in the target population, the identification features of the preset behavioral pattern are determined. Identifying these features allows for more accurate matching of user groups with the preset behavioral pattern, improving the accuracy of subsequent user group segmentation. Moreover, all of the above data processing methods are automated, improving data processing efficiency and reducing the subjectivity and bias of human judgment.

[0111] Furthermore, the method provided in this specification does not process all historical behavior data. Instead, it randomly selects the behavior data of M users from the historical behavior data, and then filters the behavior data of these M users to obtain a key behavior set. A subset of behaviors from this key behavior set is then extracted to obtain M behavior sequences in the initial population. This significantly improves the processing efficiency in subsequent data processing and reduces performance consumption. During the breeding process, this specification uses multiple breeding methods, including selection (determining the first set of behavior sequences), crossover (obtaining evolutionary behavior sequences through crossover operations), and mutation (mutating behavior sequences), to breed the initial population for multiple generations, ultimately obtaining target behavior sequences that effectively represent the preset behavior pattern. Finally, based on the target behavior sequences, the recognition features corresponding to the preset behavior pattern are determined, significantly improving the accuracy of the recognition features, the generalization ability of the model, and the accuracy of user behavior prediction. Simultaneously, it improves data processing efficiency and reduces human error.

[0112] This specification, in another aspect, provides a computer-readable non-transitory storage medium storing at least one instruction set of executable instructions for data processing. When the at least one instruction set is executed by a processor, it instructs the processor to implement the steps of the data processing method P300 of this specification. In some possible embodiments, various aspects of this specification may also be implemented as a program product comprising program code. When the program product is run on a system, the program code causes the data processing system to perform the steps of the method P300 described in this specification. The program product for implementing the above method may employ a portable compact disc read-only memory (CD-ROM) containing program code and may run on a data processing system. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the foregoing. Program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar programming languages.

[0113] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0114] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure may be presented by way of example only and may not be restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.

[0115] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.

[0116] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.

[0117] Every patent, patent application, publication of a patent application, and other material, such as articles, books, specifications, publications, documents, and literature (excluding any related historical examination documents), cited in this disclosure is incorporated herein for all purposes, including, for example, in the specification and claims of this disclosure. However, in the event of any inconsistency or conflict between the descriptions, definitions, and / or terms used in the foregoing and those used in this disclosure, the descriptions, definitions, and / or terms used in this disclosure shall prevail.

[0118] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.

Claims

1. A data processing method, comprising: Obtain historical behavioral data corresponding to the target user group, wherein the target user group includes multiple users known to have preset behavioral patterns; Based on the historical behavior data, M initial behavior sequences are generated, wherein each initial behavior sequence represents multiple behaviors performed by a user in the target user group in chronological order, and M is an integer greater than 1; The M initial behavior sequences are used as the initial population. The initial population is reproduced for multiple generations by simulating the population evolution process to obtain the target population. The target population includes N target behavior sequences. The N target behavior sequences have a higher representation ability of the preset behavior pattern than the M initial behavior sequences. N is an integer greater than 1. The i-th generation reproduction process in the multi-generational reproduction includes: A first set of behavioral sequences is formed by selecting a subset of behavioral sequences that meet the evolutionary conditions from the current population. The evolutionary conditions are related to the ability to represent the preset behavioral patterns. Perform a crossover operation on different behavior sequences in the first behavior sequence set to obtain multiple evolutionary behavior sequences, and add the multiple evolutionary behavior sequences to the first behavior sequence set to obtain a second behavior sequence set; The population after the i-th generation of reproduction is generated based on the second set of behavioral sequences, where i is any natural number. When i=1, the current population is the initial population, and when i>1, the current population is the population after the (i-1)-th generation of reproduction. Based on the target behavior sequences in the target population, the identification features corresponding to the preset behavior pattern are determined.

2. The method according to claim 1, wherein, The evolutionary behavior sequence includes multiple sub-sequences, wherein the multiple sub-sequences are respectively derived from at least two different behavior sequences in the first behavior sequence set.

3. The method according to claim 1, wherein, The generation of the i-th generation population based on the second set of behavioral sequences includes: Mutation operations are performed on a subset of the behavior sequences in the second set of behavior sequences to obtain multiple mutated behavior sequences; and The plurality of mutated behavior sequences, as well as other behavior sequences in the second set of behavior sequences that have not undergone the mutation operation, are used as the population after the i-th generation of reproduction.

4. The method according to claim 3, wherein, The mutation operation includes one or more of the following: A preset number of behaviors are randomly added to the behavior sequence; Randomly delete a preset number of behaviors from the behavior sequence; or, Randomly replace some behaviors in the behavior sequence with other behaviors.

5. The method according to claim 1, wherein, The step of selecting a subset of behavioral sequences that satisfy evolutionary conditions from the current population to form the first set of behavioral sequences includes: Obtain a fitness function corresponding to the preset behavior pattern. The fitness function uses the behavior sequence as the independent variable and fitness as the dependent variable. The fitness obtained based on the fitness function represents the degree of representation ability of a behavior sequence to the preset behavior pattern. Based on the fitness function, the fitness of each behavioral sequence in the current population is determined; Based on the fitness of each behavioral sequence in the current population, select a subset of behavioral sequences that meet the evolutionary conditions from the current population to form the first set of behavioral sequences.

6. The method according to claim 5, wherein, The fitness function includes a penalty term, which is configured to adjust the fitness of a behavior sequence based on its feasibility, and satisfies the following constraints: When the feasibility of the first action sequence is lower than that of the second action sequence, the adjusted fitness corresponding to the first action sequence is lower than that corresponding to the second action sequence.

7. The method according to claim 5, wherein, The step of selecting a subset of behavioral sequences that satisfy evolutionary conditions from the current population based on the fitness of each behavioral sequence in the current population to form the first set of behavioral sequences includes: Based on the fitness of each behavioral sequence in the current population, determine the mean fitness of the current population. The selection probability of each behavior sequence is determined based on the ratio of the mean fitness to the fitness of each behavior sequence in the current population; and From the current population, select the behavior sequences whose selection probability is greater than a preset probability threshold to form the first behavior sequence set.

8. The method according to claim 1, wherein, The termination conditions for the multigenerational reproduction include at least one of the following: The number of iterations that have been propagated is greater than or equal to the preset number of iterations, or The ability of the behavioral sequences contained in the current population to represent the preset behavioral pattern reaches the convergence condition.

9. The method according to claim 1, wherein, The generation of M initial behavior sequences based on the historical behavior data includes: Randomly select the behavior data of M users from the historical behavior data; Obtain a set of key behaviors corresponding to the preset behavior pattern, wherein the set of key behaviors includes multiple key behaviors related to the preset behavior pattern; and For each of the M users, select multiple behaviors from the user's behavior data that are located in the key behavior set, and generate an initial behavior sequence based on the multiple behaviors.

10. The method according to claim 1, wherein, The method further includes: Users with the preset behavioral patterns are extracted from the user set based on the identified features.

11. A data processing system, comprising: At least one storage medium storing at least one instruction set for performing data processing methods; as well as At least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set during operation and performs a data processing method as described in any one of claims 1-10 according to the instructions of the at least one instruction set.

Citation Information

Patent Citations

  • Abnormal user detection method based on fuzzy sequential association pattern

    CN105262715A

  • Method and device for determining parameters of sorting model in recommendation system

    CN110457545A