A cash flow data analysis method based on artificial intelligence
By building a priori tag library and periodically parsing tag pools, and analyzing the tag association relationships in cash flow data, the problem of inconsistent tags in user portraits is solved, and the accuracy and adaptability of information push are achieved.
Patent Information
- Application Number
- CN202411731175.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-29
AI Technical Summary
When constructing user portraits or labels, existing technologies lead to inconsistent data responses due to the differences and sporadic nature of different users' behavioral habits, introduce labels with poor data representation, and cause inaccurate information push.
By acquiring cash flow data, building a priori tag library, analyzing the correlation between various tags, periodically parsing the tag pool, verifying the validity of the tag pool, updating the tag pool, generating a verification sub-tag pool for verification push, and eliminating invalid tags, the accuracy of the tag pool is ensured.
It improves the accuracy of tags in the tag pool, ensures the accuracy of information push, adapts to the individual differences and sporadic nature of user behavior, and ensures the effectiveness of information push.
Smart Images

Figure CN119693050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular to a cash flow data analysis method based on artificial intelligence. Background Art
[0002] Cash flow data includes key information such as the transaction object, transaction amount, and transaction goods during the user's transaction process. Collecting cash flow data under the conditions of big data acquisition permission has broad application value. For example, through artificial intelligence, the collected data can be analyzed to obtain user needs or set labels for users, so as to facilitate the subsequent targeted push of relevant recommended information to users.
[0003] Chinese patent publication number CN109598540A discloses a method and system for precise advertising push. The method comprises: parsing commodity transaction data pre-stored on a blockchain platform to identify associated transaction commodity information and transaction user accounts; comparing the transaction commodity information with the pre-stored push information for similarity; and, if the similarity exceeds a threshold, pushing the push information to the transaction user account associated with the transaction commodity information. The method and system of the present invention record commodity transaction information based on the effectiveness of commodity transaction smart contracts on the blockchain platform and publish it to the blockchain platform, ensuring the authenticity of commodity transactions. The method and system select advertisement push targets based on the commodity transaction information, achieving efficient advertisement push.
[0004] However, the prior art still has the following problems:
[0005] In reality, when building user portraits or labels for the user side, the differences in user behavior habits lead to inconsistent responses at the data level. In addition, due to the sporadic nature of user behavior, some labels with poor data representation will be introduced, resulting in inaccurate information push. Summary of the Invention
[0006] To this end, the present invention provides an artificial intelligence-based cash flow data analysis method to overcome the problem in the prior art that when constructing user portraits or labels for the user side, the differences in the behavioral habits of different users lead to inconsistent responses at the data level, and due to the sporadic nature of user behavior, some labels with poor data representation are introduced, resulting in inaccurate information push.
[0007] To achieve the above objectives, the present invention provides a cash flow data analysis method based on artificial intelligence, comprising:
[0008] Obtain access permissions for cash flow data of several user terminals, generate labels based on the cash flow data within a historical period, and build a priori label library;
[0009] Determine the association between various tags based on the proportion of each tag in the total tags corresponding to the user terminal in the prior tag library and the conditional probability that each tag belongs to the same user terminal;
[0010] Construct a tag pool for the user end, generate tags in real time based on the cash flow data updated by the user end, and store them in the tag pool. Periodically perform correlation analysis on the tag pool, including determining the tag correlation coefficient based on the correlation relationship between various tags to verify whether the validity of the tag pool meets the standards.
[0011] Updating the user terminal's tag pool based on the verification result of the tag pool validity includes:
[0012] Based on the association between each tag, call the tag to generate a verification sub-tag pool, generate push information based on the verification sub-tag pool, perform verification push and collect the user's response data to the push information, and remove the tag in the corresponding verification sub-tag pool based on the response data;
[0013] Or, maintain the tags in the current tag pool;
[0014] Periodically generate push information based on the user's tag pool and push it to the user.
[0015] Furthermore, the process of generating labels based on cash flow data includes:
[0016] Pre-set several data features required to be extracted from cash flow data;
[0017] Extracting several data features from the cash flow data, and generating corresponding labels according to each of the data features;
[0018] There is a one-to-one correspondence between each of the data features and the label.
[0019] Furthermore, determining whether there is a correlation between various tags includes:
[0020] Determine the average probability that each tag group belongs to the same user terminal in different periods;
[0021] Determine the proportion of each tag group in all tags corresponding to the user terminal to which it belongs;
[0022] If the association condition is met, it is determined that there is an association relationship between the corresponding category tags in the tag group and the degree of association is determined;
[0023] The tag group includes two types of tags, the association condition is that the average probability is greater than a predetermined probability threshold and the proportion is greater than a predetermined proportion threshold, and the association degree is the average probability corresponding to the tag group.
[0024] Furthermore, the process of determining the label correlation coefficient based on the correlation relationship between various labels includes:
[0025] Determine the ratio of the number of tags with associated relationships in the tag pool to the total number of tags;
[0026] Determine the mean correlation between various tags that have correlation relationships;
[0027] The ratio of the proportion to a preset correlation proportion threshold is used as a first correlation feature;
[0028] The ratio of the correlation degree mean to a preset correlation degree threshold is used as a second correlation feature;
[0029] The sum of the first association feature and the second association feature is determined as a tag association coefficient.
[0030] Furthermore, the process of verifying whether the validity of the tag pool meets the standards includes:
[0031] If the tag association coefficient is greater than or equal to a preset tag association threshold, then verifying that the validity of the tag pool meets the standard;
[0032] If the tag association coefficient is less than a preset tag association threshold, it is verified that the validity of the tag pool does not meet the standard.
[0033] Further, updating the tag pool of the user terminal based on the verification result of the validity of the tag pool includes:
[0034] If the validity of the verification tag pool does not meet the standards, the tag is called to generate several verification sub-tag pools based on the association relationship between the tags. Push information is generated based on the verification sub-tag pool to perform verification push and collect the user end's response data to the push information. Based on the response data, the tag in the corresponding verification sub-tag pool is removed;
[0035] If the validation of the label pool meets the criteria, the labels in the current label pool are maintained.
[0036] Furthermore, the process of generating several verification sub-tag pools includes:
[0037] Identify irrelevant and potentially relevant tags;
[0038] The unrelated tags and the potentially related tags are stored to generate a verification sub-tag pool;
[0039] The unrelated tag satisfies the requirement that it has no association relationship with any of the remaining tags, and the potential associated tag satisfies the requirement that the degree of association between it and any of the remaining tags is greater than a preset potential association threshold.
[0040] Furthermore, the process of generating push information based on the verification sub-tag pool and performing verification push includes:
[0041] Calling the unrelated tags and potential related tags in the verification sub-tag pool;
[0042] Generate corresponding push information based on unrelated tags and potentially related tags;
[0043] Push the generated push information to the user end;
[0044] Among them, the corresponding relationship between each tag and the required push information is preset.
[0045] Furthermore, the process of obtaining the user-side response data includes:
[0046] Determine the user's response data to the push information during multiple push verification processes, including whether it was clicked and the browsing time;
[0047] Determine the number of valid clicks for each type of push information based on browsing time, and calculate the variance and mean of the number of valid clicks for each type of push information;
[0048] Among them, if the browsing time of a single click on the pushed information is greater than the predetermined browsing time threshold, it is determined to be a valid click.
[0049] Furthermore, removing tags from the corresponding verification subtag pool based on the response data includes:
[0050] If the benchmark elimination condition is met, all tags in the verification sub-tag pool are determined to be eliminated;
[0051] If the benchmark rejection condition is not met, it is determined to reject the invalid verification tag in the verification sub-tag pool;
[0052] The benchmark elimination condition includes that the variance of the number of valid clicks is greater than a predetermined variance threshold and the mean number of valid clicks is less than a predetermined click benchmark threshold;
[0053] The invalid verification tag satisfies the requirement that the number of valid clicks on the pushed information is less than a predetermined click reference threshold.
[0054] Compared with the existing technology, the present invention generates tags based on cash flow data in a historical period, builds a priori tag library, analyzes the correlation between various tags, builds a tag pool for the user end, and periodically performs correlation analysis on the tag pool, including determining the tag correlation coefficient based on the correlation between various tags, verifying whether the validity of the tag pool meets the standards, and subsequently adaptively updating the tag pool of the user end. When the validity of the tag pool does not meet the standards, a verification sub-tag pool is generated, push information is generated based on the verification sub-tag pool, and tags in the sub-tag pool are removed based on the response data of the push information. The present invention can improve the accuracy of tags in the tag pool, thereby ensuring the accuracy of push information.
[0055] In particular, the present invention considers constructing a priori tag library to determine the association relationship between various tags. In actual situations, the cash flow data of the user terminal has data characteristics, and corresponding tags can be generated. Due to individual differences, the cash flow data of different user terminals are different. However, for a single user terminal, the corresponding data characteristics between its cash flow data usually have potential associations. Therefore, the present invention takes the above characteristics into consideration and pre-constructs the association relationship between various tags through the attribution relationship of tags in a specific priori tag library, providing data support for the subsequent verification of the user-specific tag pool to ensure the validity of the tag pool, and then ensure the accuracy of subsequent information push.
[0056] In particular, the present invention considers periodically performing association analysis on the label pool. In actual situations, due to the unpredictability of individual behavior, when collecting cash flow data to determine labels, some cash flow data may be sporadic data, and the labels generated are poorly representative of the data on the user side. As the labels in the label pool accumulate, the labels generated by sporadic data are also accumulating, making the effectiveness of the label pool poor. Therefore, the present invention considers performing association analysis on the label pool and calculating the label association coefficient. The label association coefficient is calculated from data of two dimensions and comprehensively represents the association of labels in the label pool. If the label association coefficient is low, it means that too many labels with poor data representation are introduced, and the effectiveness of the label pool can be evaluated in a timely manner to ensure the accuracy of subsequent information push.
[0057] In particular, when the validity of the tag pool does not meet the standards, the present invention specifically generates a verification sub-tag pool, generates push information based on the verification sub-tag pool for verification push, and selectively eliminates some tags. In actual situations, when the tag pool does not meet the standards, there are many tags with poor data representativeness, so these tags are generated into a verification sub-tag pool. At the same time, considering that there may be potential associations among these tags, push information is generated based on the entire verification sub-tag pool, and verification push is performed. The tags in the verification sub-tag pool are verified based on the response data, and then the tag pool of the user end can be updated in time to ensure the validity of the tag pool and the accuracy of subsequent information push. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A step diagram of a cash flow data analysis method based on artificial intelligence according to an embodiment of the invention;
[0059] Figure 2 A logical determination diagram for determining whether there is an association relationship between various tags in an embodiment of the invention;
[0060] Figure 3 A logical decision diagram for verifying whether the validity of a tag pool conforms to a standard according to an embodiment of the invention;
[0061] Figure 4 This is a logical decision diagram for updating the tag pool of the user terminal based on the verification result of the validity of the tag pool according to an embodiment of the invention. DETAILED DESCRIPTION
[0062] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0063] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0064] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0065] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0066] See also Figures 1 to 4 As shown, Figure 1 This is a step diagram of the cash flow data analysis method based on artificial intelligence according to an embodiment of the invention. Figure 2 This is a logical decision diagram for determining whether there is an association relationship between various tags in the embodiment of the invention. Figure 3 The logic decision diagram for verifying whether the validity of the tag pool of the embodiment of the invention meets the standard is as follows: Figure 4 In order to update the logic decision diagram of the tag pool of the user terminal based on the verification result of the validity of the tag pool according to the embodiment of the invention, the cash flow data analysis method based on artificial intelligence provided by the embodiment of the invention includes:
[0067] Step S1: obtaining access permissions for cash flow data of several user terminals, generating labels based on the cash flow data in a historical period, and building a priori label library;
[0068] Step S2: determining the association between various tags based on the proportion of each tag in the total tags corresponding to the user terminal in the prior tag library and the conditional probability that each tag belongs to the same user terminal;
[0069] Step S3: Construct a tag pool for the user terminal, generate tags in real time based on the cash flow data updated by the user terminal, store them in the tag pool, and periodically perform association analysis on the tag pool, including determining the tag association coefficient based on the association relationship between various tags to verify whether the validity of the tag pool meets the standards;
[0070] Step S4, based on the verification result of the validity of the tag pool, updates the tag pool of the user terminal, including:
[0071] Based on the association between each tag, call the tag to generate a verification sub-tag pool, generate push information based on the verification sub-tag pool, perform verification push and collect the user's response data to the push information, and remove the tag in the corresponding verification sub-tag pool based on the response data;
[0072] Or, maintain the tags in the current tag pool;
[0073] Step S5: periodically generate push information based on the tag pool of the user terminal and push it to the user terminal.
[0074] Specifically, cash flow data includes the transaction objects and transaction commodities in the user's transaction process. It can be understood that cash flow data is multidimensional, in which both the transaction objects and transaction commodities can be used as data features.
[0075] Specifically, the process of generating labels based on cash flow data includes:
[0076] Pre-set several data features required to be extracted from cash flow data;
[0077] Extracting several data features from the cash flow data, and generating corresponding labels according to each of the data features;
[0078] There is a one-to-one correspondence between each of the data features and the label.
[0079] In some possible implementations, taking trading goods as data features as an example, corresponding labels can be generated based on the categories of trading goods. For example, trading goods can be divided into furniture, daily necessities, books, hardware, etc. Of course, the categories can be further subdivided, and then the categories can be used as labels. Of course, it can also be any other form, such as dividing according to the use areas of trading goods, for example, not limited to field categories, mining categories, automotive categories, etc. Similarly, categories can be further subdivided. It can be understood that the setting of labels is based on the needs of technical personnel in this field and can be changed, which will not be repeated here.
[0080] In some possible implementations, taking the transaction object as a data feature as an example, it can be divided according to the category of the transaction object, for example, official stores, private stores, personal transactions, etc. Similarly, the label setting is based on the needs of technical personnel in this field and can be changed, which will not be repeated here.
[0081] Specifically, determining whether there is a correlation between various tags includes:
[0082] Determine the average probability of each tag group belonging to the same user terminal in different periods. It can be understood that the average probability of a tag group belonging to the same user terminal can represent the probability of a single user terminal generating data features corresponding to the tags in the tag group in the data dimension, and thus indirectly represent the possible associations between data features;
[0083] Determine the proportion of each tag group in all tags corresponding to the user terminal to which it belongs. It can be understood that the purpose of calculating the proportion is to filter out the tag group with a higher proportion and ensure the effectiveness of the tag group relative to the user terminal;
[0084] If the association condition is met, it is determined that there is an association relationship between the corresponding category tags in the tag group and the degree of association is determined;
[0085] The tag group includes two types of tags, the association condition is that the average probability is greater than a predetermined probability threshold and the proportion is greater than a predetermined proportion threshold, and the association degree is the average probability corresponding to the tag group.
[0086] Specifically, the average probability that a tag group belongs to the same user terminal is the probability that various tags included in the tag group co-appear in the same user terminal, which will not be described in detail.
[0087] Specifically, the predetermined probability threshold and the predetermined proportion threshold are pre-set. The cash flow data of several user terminals during the experimental period are collected in advance to determine the average probability that the tag group belongs to the same user terminal, as well as the average proportion of each tag group of a single user terminal. The predetermined probability threshold is set as the product of the average probability and the first precision coefficient, and the predetermined proportion threshold is set as the product of the average proportion and the second precision coefficient. The first precision coefficient is selected within the interval [1.25, 1.5], and the second precision coefficient is selected within the interval [1.15, 1.3].
[0088] It is understandable that in order to ensure sample accuracy, the experimental period can be longer than the historical period corresponding to the construction of the prior label library. The historical period can be selected within one to two months, and the experimental period can be selected within three to five months.
[0089] Specifically, the present invention considers constructing a priori label library to determine the association relationship between various labels. In actual situations, the cash flow data of the user end has data characteristics, and corresponding labels can be generated. Due to individual differences, the cash flow data of different user ends are different. However, for a single user end, the corresponding data characteristics between its cash flow data usually have potential associations. Therefore, the present invention takes the above characteristics into consideration and pre-constructs the association relationship between various labels through the attribution relationship of labels in a specific priori label library, providing data support for the subsequent verification of the user-specific label pool to ensure the validity of the label pool, and then ensure the accuracy of subsequent information push.
[0090] Specifically, the process of determining the label correlation coefficient based on the correlation between various labels includes:
[0091] Determine the ratio of the number of tags with associated relationships in the tag pool to the total number of tags;
[0092] Determine the mean correlation between various tags that have correlation relationships;
[0093] The ratio of the proportion to a preset correlation proportion threshold is used as a first correlation feature;
[0094] The ratio of the correlation degree mean to a preset correlation degree threshold is used as a second correlation feature;
[0095] The sum of the first association feature and the second association feature is determined as a tag association coefficient.
[0096] Specifically, the preset association ratio threshold is pre-set, and a tag pool for several user terminals is pre-built. The average ratio of the number of tags with associated relationships in the tag pool corresponding to each user terminal to the total number of tags is solved, and the preset association ratio threshold is set to between 1.3 and 1.6 times the average ratio.
[0097] Specifically, the correlation threshold is pre-set. Cash flow data of several user terminals within the experimental period are collected in advance, tags with correlation relationships are determined, the average correlation of each tag is solved, and the correlation threshold is set to between 1.25 and 1.5 times the average correlation.
[0098] Specifically, the process of verifying whether the validity of the tag pool meets the standards includes:
[0099] If the tag association coefficient is greater than or equal to a preset tag association threshold, then verifying that the validity of the tag pool meets the standard;
[0100] If the tag association coefficient is less than a preset tag association threshold, it is verified that the validity of the tag pool does not meet the standard.
[0101] Specifically, the label association threshold is selected in the interval [2.15, 2.3].
[0102] Specifically, the present invention considers periodically performing association analysis on the label pool. In actual situations, due to the unpredictability of individual behavior, when collecting cash flow data to determine labels, some cash flow data may be sporadic data, and the labels generated are poorly representative of the data on the user side. As the labels in the label pool accumulate, the labels generated by sporadic data are also accumulating, making the effectiveness of the label pool poor. Therefore, the present invention considers performing association analysis on the label pool and calculating the label association coefficient. The label association coefficient is calculated from data of two dimensions and comprehensively characterizes the association of labels in the label pool. If the label association coefficient is low, it means that too many labels with poor data representation are introduced, and the effectiveness of the label pool can be evaluated in a timely manner to ensure the accuracy of subsequent information push.
[0103] Specifically, updating the tag pool of the user terminal based on the verification result of the validity of the tag pool includes:
[0104] If the validity of the verification tag pool does not meet the standards, the tag is called to generate several verification sub-tag pools based on the association relationship between the tags. Push information is generated based on the verification sub-tag pool to perform verification push and collect the user end's response data to the push information. Based on the response data, the tag in the corresponding verification sub-tag pool is removed;
[0105] If the validation of the label pool meets the criteria, the labels in the current label pool are maintained.
[0106] Specifically, the process of generating several verification sub-tag pools includes:
[0107] Identify irrelevant and potentially relevant tags;
[0108] The unrelated tags and the potentially related tags are stored to generate a verification sub-tag pool;
[0109] The unrelated tag satisfies the requirement that it has no association relationship with any of the remaining tags, and the potential associated tag satisfies the requirement that the degree of association between it and any of the remaining tags is greater than a preset potential association threshold.
[0110] Specifically, the potential association threshold is pre-set to 0.85 times the association threshold. It can be understood that the purpose of setting the potential association threshold is to determine tags that have a potential association but whose association with other tags does not reach the association threshold.
[0111] Specifically, the process of generating push information based on the verification sub-tag pool for verification push includes:
[0112] Calling the unrelated tags and potential related tags in the verification sub-tag pool;
[0113] Generate corresponding push information based on unrelated tags and potentially related tags;
[0114] Push the generated push information to the user end;
[0115] Among them, the corresponding relationship between each tag and the required push information is preset.
[0116] Specifically, in some possible implementations, push information takes advertising as an example. For example, if the label is furniture, then furniture advertising information will be pushed. It can be understood that the label and the push information are not in a one-to-one correspondence. A single label can correspond to multiple push information. Those skilled in the art can set it according to their needs, and this will not be repeated.
[0117] Specifically, the process of obtaining the user-side response data includes:
[0118] Determine the user's response data to the push information during multiple push verification processes, including whether it was clicked and the browsing time;
[0119] Determine the number of valid clicks for each type of push information based on browsing time, and calculate the variance and mean of the number of valid clicks for each type of push information;
[0120] Among them, if the browsing time of a single click on the pushed information is greater than the predetermined browsing time threshold, it is determined to be a valid click.
[0121] The browsing time can be determined as the browsing time based on the time spent on the corresponding push information page after clicking the push information, which will not be described in detail.
[0122] Specifically, the browsing time threshold is intended to indicate that the user terminal browses the pushed information, and those skilled in the art may select it within the range of [2s, 5s].
[0123] Specifically, based on the response data, the tags in the corresponding verification sub-tag pool are removed,
[0124] If the benchmark elimination condition is met, all tags in the verification sub-tag pool are determined to be eliminated;
[0125] If the benchmark rejection condition is not met, it is determined to reject the invalid verification tag in the verification sub-tag pool;
[0126] The benchmark elimination condition includes that the variance of the number of valid clicks is greater than a predetermined variance threshold and the mean number of valid clicks is less than a predetermined click benchmark threshold;
[0127] It is understandable that when the benchmark elimination conditions are met, the user's tendency towards various types of push information is relatively discrete, and the correlation between the labels in the verification sub-label pool is weak. They may all be labels generated by relatively occasional cash flow data on the user side, and therefore are eliminated accordingly.
[0128] In the implementation, the predetermined variance threshold is pre-set, wherein verification push is performed for user terminals in a valid tag pool for several tag pools, including generating push information based on the tags in the valid tag pool, recording response data, determining the mean variance of the number of valid clicks for each type of push information, and setting the predetermined variance threshold to between 1.25 and 1.5 times the mean variance of the number of valid clicks.
[0129] In practice, the click count benchmark threshold is determined based on the number of push verifications, and is usually set between 0.3 and 0.5 times the number of push verifications.
[0130] The invalid verification tag satisfies the requirement that the number of valid clicks on the pushed information is less than a predetermined click reference threshold.
[0131] It is understandable that if the number of valid clicks on the push information is less than the predetermined click threshold, it indicates that the user terminal has a weak tendency to push this type of information, and therefore the corresponding label is removed.
[0132] Specifically, when the validity of the label pool does not meet the standards, the present invention specifically generates a verification sub-label pool, generates push information based on the verification sub-label pool for verification push, so as to selectively eliminate some labels. In actual situations, when the label pool does not meet the standards, there are many labels with poor data representativeness, so these labels are generated into a verification sub-label pool. At the same time, considering that there may be potential associations among these labels, push information is generated based on the entire verification sub-label pool, and verification push is performed. The labels in the verification sub-label pool are supported based on the response data, and thus the label pool of the user end can be updated in a timely manner to ensure the validity of the label pool and the accuracy of subsequent information push.
[0133] If the artificial intelligence-based cash flow data analysis method of the present invention is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0134] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A cash flow data analysis method based on artificial intelligence, characterized in that: include: Obtain access permissions for cash flow data of several user terminals, generate labels based on the cash flow data within a historical period, and build a priori label library; Determine the association between various tags based on the proportion of each tag in the total tags corresponding to the user terminal in the prior tag library and the conditional probability that each tag belongs to the same user terminal; Construct a tag pool for the user end, generate tags in real time based on the cash flow data updated by the user end, and store them in the tag pool. Periodically perform correlation analysis on the tag pool, including determining the tag correlation coefficient based on the correlation relationship between various tags to verify whether the validity of the tag pool meets the standards. Based on the verification result of the validity of the tag pool, the tag pool of the user terminal is updated. include, Based on the association between each tag, call the tag to generate a verification sub-tag pool, generate push information based on the verification sub-tag pool, perform verification push and collect the user's response data to the push information, and remove the tag in the corresponding verification sub-tag pool based on the response data; Or, maintain the tags in the current tag pool; Periodically generate push information based on the user's tag pool and push it to the user.
2. The cash flow data analysis method based on artificial intelligence according to claim 1 is characterized in that: The process of generating labels based on cash flow data includes: Pre-set several data features required to be extracted from cash flow data; Extracting several data features from the cash flow data, and generating corresponding labels according to each of the data features; There is a one-to-one correspondence between each of the data features and the label.
3. The cash flow data analysis method based on artificial intelligence according to claim 2 is characterized in that: Determine whether there is a correlation between various tags, including: Determine the average probability that each tag group belongs to the same user terminal in different periods; Determine the proportion of each tag group in all tags corresponding to the user terminal to which it belongs; If the association condition is met, it is determined that there is an association relationship between the corresponding category tags in the tag group and the degree of association is determined; The tag group includes two types of tags, the association condition is that the average probability is greater than a predetermined probability threshold and the proportion is greater than a predetermined proportion threshold, and the association degree is the average probability corresponding to the tag group.
4. The cash flow data analysis method based on artificial intelligence according to claim 3 is characterized in that: The process of determining the label correlation coefficient based on the correlation between various labels includes: Determine the ratio of the number of tags with associated relationships in the tag pool to the total number of tags; Determine the mean correlation between various tags that have correlation relationships; The ratio of the proportion to a preset correlation proportion threshold is used as a first correlation feature; The ratio of the correlation degree mean to a preset correlation degree threshold is used as a second correlation feature; The sum of the first association feature and the second association feature is determined as a tag association coefficient.
5. The cash flow data analysis method based on artificial intelligence according to claim 1 is characterized in that: The process of verifying whether the validity of the tag pool meets the standards includes: If the tag association coefficient is greater than or equal to a preset tag association threshold, then verifying that the validity of the tag pool meets the standard; If the tag association coefficient is less than a preset tag association threshold, it is verified that the validity of the tag pool does not meet the standard.
6. The cash flow data analysis method based on artificial intelligence according to claim 1 is characterized in that: Updating the user terminal's tag pool based on the verification result of the tag pool validity includes: If the validity of the verification tag pool does not meet the standards, the tag is called to generate several verification sub-tag pools based on the association relationship between the tags. Push information is generated based on the verification sub-tag pool to perform verification push and collect the user end's response data to the push information. Based on the response data, the tag in the corresponding verification sub-tag pool is removed; If the validation of the label pool meets the criteria, the labels in the current label pool are maintained.
7. The cash flow data analysis method based on artificial intelligence according to claim 1 is characterized in that: The process of generating several verification sub-tag pools includes: Identify irrelevant and potentially relevant tags; The unrelated tags and the potentially related tags are stored to generate a verification sub-tag pool; The unrelated tag satisfies the requirement that it has no association relationship with any of the remaining tags, and the potential associated tag satisfies the requirement that the degree of association between it and any of the remaining tags is greater than a preset potential association threshold.
8. The cash flow data analysis method based on artificial intelligence according to claim 7 is characterized in that: The process of generating push information based on the verification sub-tag pool for verification push includes: Calling the unrelated tags and potential related tags in the verification sub-tag pool; Generate corresponding push information based on unrelated tags and potentially related tags; Push the generated push information to the user end; Among them, the corresponding relationship between each tag and the required push information is preset.
9. The cash flow data analysis method based on artificial intelligence according to claim 1 is characterized in that: The process of obtaining the user-side response data includes: Determine the user's response data to the push information during multiple push verification processes, including whether it was clicked and the browsing time; Determine the number of valid clicks for each type of push information based on browsing time, and calculate the variance and mean of the number of valid clicks for each type of push information; Among them, if the browsing time of a single click on the pushed information is greater than the predetermined browsing time threshold, it is determined to be a valid click.
10. The cash flow data analysis method based on artificial intelligence according to claim 1, characterized in that: Eliminating tags in the corresponding verification sub-tag pool based on the response data includes: If the benchmark elimination condition is met, all tags in the verification sub-tag pool are determined to be eliminated; If the benchmark rejection condition is not met, it is determined to reject the invalid verification tag in the verification sub-tag pool; The benchmark elimination condition includes that the variance of the number of valid clicks is greater than a predetermined variance threshold and the mean number of valid clicks is less than a predetermined click benchmark threshold; The invalid verification tag satisfies the requirement that the number of valid clicks on the pushed information is less than a predetermined click reference threshold.
Citation Information
Patent Citations
A precise advertisement pushing method and a precise advertisement pushing system
CN109598540A
Pushing method, user label generation method and device, and equipment
CN108881339A
Commodity recommendation method and system based on big data and artificial intelligence
CN118014684A