A feature learning-based data pre-management method and system

By monitoring user session behavior and learning its features, recording session dwell time and pointer records, fitting habit feature curves, marking preference change nodes, and establishing a data pre-screening model, the problem of inaccurate screening of noisy data is solved, and the signal-to-noise ratio and accuracy of user preference judgment are improved.

CN120873532BActive Publication Date: 2025-12-09CHENGDU POLYTECHNIC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511383204.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-09
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing technologies often fail to accurately filter noisy data, leading to low accuracy and efficiency in judging user preferences, especially when dealing with multiple topics.

Method used

By monitoring user session behavior, recording session dwell and pointer recording, calculating deviation characteristics, fitting habit feature curves, marking preference change nodes, establishing a data pre-screening model, identifying noisy data, and optimizing the screening process through cross-validation and attention evaluation.

Benefits of technology

It improved the signal-to-noise ratio of user preference judgment, optimized evaluation efficiency and accuracy, avoided the problem of inaccurate filtering under multiple topics, and improved the accuracy of data push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873532B_ABST
    Figure CN120873532B_ABST
Patent Text Reader

Abstract

The application relates to the field of user data screening, and discloses a data pre-management method and system based on feature learning, which monitors and records the content browsing sessions of users, judges individual behavior features, pre-screens a large amount of session data, pre-screens noise data useless for user content preference judgment, improves the signal-to-noise ratio of a database in subsequent user content preference judgment, optimizes evaluation efficiency and accuracy, and compared with the existing noise data screening mode based on content judgment session preference, the noise screening mode based on behavior features can more accurately avoid the inaccurate screening caused by multiple content themes in a single session.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of user data screening, and particularly relates to a data pre-management method and system based on feature learning. BACKGROUND

[0002] User preference judgment and data pushing based on big data are widely used in various fields in the current Internet environment, including but not limited to content platform pushing, user-oriented screening of sales platforms, etc. Therefore, the accuracy of user preference judgment is very important for data pushing. Noise screening of a large amount of user data before preference judgment can greatly reduce noise data in preference fitting judgment, so that the preference judgment result is more accurate.

[0003] The noise data screening and cleaning in the prior art mainly includes two categories. One is to set a stay rule to screen low-quality data, that is, if the user stays on the content page for less than a preset value, it means that the content is noise data for user preference judgment. The second is to preliminarily screen the content type by using the existing user conversion label, but both of these two ways will be affected by the total amount of information content and the multi-theme of information content, resulting in a great reduction in the accuracy of noise data screening. Ultimately, only a small part of noise can be screened out, and most of the noise data is classified and judged in the preference fitting. This way not only increases the efficiency of preference judgment, but also reduces the accuracy of preference judgment, so there is a certain optimization space. SUMMARY

[0004] The present application aims to provide a data pre-management method and system based on feature learning to solve the problems raised in the background.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical scheme:

[0006] A data pre-management method based on feature learning, comprising:

[0007] Conducting session monitoring on the user to record the session behavior characteristics of the user, the session behavior characteristics representing the browsing behavior of the user on the session content, including session stay records and session pointer records;

[0008] Calculating the deviation characteristics based on the online average stay of each session and the session stay records to generate the session stay characteristics of the current session records, the online average stay being used to represent the average browsing stay duration of other users, the deviation characteristics calculation being used to represent the process of calculating the time difference value and the average browsing stay duration proportion, and the session stay characteristics having directionality;

[0009] The several session stay features are sequentially counted to perform curve fitting, the habit feature curve of the user is acquired, and the rising turning point of the habit feature curve is judged to mark as the preference change node of the user.

[0010] A data pre-screening model is established based on the preference change node, the session behavior feature is evaluated according to the data pre-screening model to judge whether it is noise data, and the data pre-screening model comprises a deviation feature calculation model and a comparison and judgment program based on the preference change node.

[0011] As a further scheme of the application, the step of sequentially counting the several session stay features to perform curve fitting, acquiring the habit feature curve of the user, and judging the rising turning point of the habit feature curve comprises:

[0012] The data type of the session corresponding to the session stay feature is judged and distinguished to acquire a plurality of feature data sets;

[0013] The several session stay features in each feature data set are sequentially counted respectively to establish a scatter plot, and adjacent scatters are distributed at equal intervals on the horizontal coordinate.

[0014] The scatter plot is curve fitted, the change curvature at each scatter is calculated, the curve trend is judged based on the curvature, and a flat interval with a horizontal or preset low inclination angle is acquired.

[0015] The left end point of the flat interval is set as the preference change node, and if the number of flat intervals is multiple, the left end points of the first flat interval are removed to set multiple preference change nodes corresponding to different preference levels.

[0016] As a further scheme of the application, the step of sequentially counting the several session stay features to perform curve fitting, acquiring the habit feature curve of the user, and judging the rising turning point of the habit feature curve comprises:

[0017] The in-out node coordinates of the corresponding session output window are acquired, and the multiple session pointer records are plane-aligned based on the in-out node coordinates.

[0018] The several aligned session pointer records are screened, the shortest pointer record between the in-out coordinate nodes is acquired and taken as a reference pointer record.

[0019] The reference pointer record is discretely evaluated, the aggregation area of the reference pointer record is acquired, and is marked as a noise area.

[0020] If the session behavior feature is represented in the noise area, the current session is marked as noise data.

[0021] As a further scheme of the application, the step of sequentially counting the several session stay features to perform curve fitting, acquiring the habit feature curve of the user, and judging the rising turning point of the habit feature curve comprises:

[0022] The noise data result of noise area judgment is cross-judged through the noise data result of data pre-screening model judgment, if both results are represented as noise data, the noise data is finally marked, otherwise, no noise data marking is performed.

[0023] As a further scheme of the present application, the method further comprises the steps of:

[0024] If the session pointer record exceeds the noise area, the concentration of the session pointer is evaluated, and a plurality of content concentration areas are obtained, the content concentration area is used to represent the area of the high residence time proportion of the pointer in the session pointer record, and the high residence proportion is selected by area proportion sorting;

[0025] The content concentration area is marked and mapped in the session output window to mark the session content.

[0026] The embodiment of the present application aims to provide a data pre-management system based on feature learning, comprising:

[0027] The session recording module is used for monitoring the session of the user to record the session behavior characteristics of the user, the session behavior characteristics represent the browsing behavior of the user to the session content, including session residence record and session pointer record;

[0028] The feature calculation module is used for calculating the deviation feature based on the online average residence and the session residence record of each session to generate the session residence feature of the current session record, the online average residence is used to represent the average browsing residence time of other users, the deviation feature calculation is used to represent the process of calculating the time difference value and the average browsing residence time proportion, and the session residence feature has directionality;

[0029] The fitting and splitting module is used for sequentially calculating a plurality of session residence features to perform curve fitting, obtaining the habit feature curve of the user, and judging the rising turning node of the habit feature curve to mark the preference change node of the user;

[0030] The modeling and screening module is used for establishing a data pre-screening model based on the preference change node, evaluating the session behavior characteristics according to the data pre-screening model to judge whether it is noise data, and the data pre-screening model includes a deviation feature calculation model and a comparison judgment program based on the preference change node.

[0031] As a further scheme of the present application, the fitting and splitting module comprises:

[0032] The content classification unit is used for judging and distinguishing the data type of the session corresponding to the session residence feature to obtain a plurality of feature data sets;

[0033] A statistical mapping unit is configured to sequentially statistically analyze a plurality of session stay features in each feature data set to establish a scatter plot image, and adjacent scatters are distributed at equal intervals on the horizontal coordinate.

[0034] A curve fitting unit is configured to perform curve fitting on the scatter plot image, calculate the change curvature at each scatter, determine the curve trend based on the curvature, and obtain a flat interval with a horizontal or preset low inclination angle as the curve trend representation.

[0035] A node setting unit is configured to set the left end point of the flat interval as a preference change node, and if there are a plurality of flat intervals, the left end points of the plurality of flat intervals except the first flat interval are set as a plurality of preference change nodes corresponding to different preference levels.

[0036] As a further scheme of the present application, a pointer pre-selection module is further included, which specifically includes:

[0037] A record alignment unit is configured to obtain the in-out node coordinates corresponding to the session output window, and perform planar alignment on a plurality of session pointer records based on the in-out node coordinates.

[0038] A reference selection unit is configured to filter the plurality of aligned session pointer records, obtain the shortest pointer record between the in-out coordinate nodes, and take the shortest pointer record as a reference pointer record.

[0039] A discrete selection unit is configured to perform discrete evaluation based on the reference pointer record, obtain the aggregation area of the reference pointer record, and mark the aggregation area as a noise area.

[0040] A noise marking unit is configured to mark the current session as noise data if the session behavior feature represents that the session behavior feature is in the noise area.

[0041] As a further scheme of the present application, a cross-validation unit is further included.

[0042] The cross-validation unit is configured to cross-validate the noise data result of the noise area judgment based on the noise data result of the data pre-filtering model judgment, and if both results represent noise data, the noise data is finally marked as noise data, otherwise, no noise data marking is performed.

[0043] As a further scheme of the present application, a focus judgment module is further included, which specifically includes:

[0044] A pointer focus judgment unit is configured to perform focus evaluation on the session pointer if the session pointer record is out of the noise area, obtain a plurality of content focus areas, the content focus area is used to represent a high stay duration proportion area of the pointer in the session pointer record, and the high stay proportion is selected by area proportion sorting.

[0045] A content focus mapping unit is configured to mark the content focus area and map the marked content focus area to a conversation output window to mark the conversation content.

[0046] Compared with the prior art, the present application has the beneficial effects that: by monitoring and recording the content browsing session of the user, the individual behavior characteristics are judged, a large amount of conversation data is pre-screened, noise data useless for judging the content preference of the user is pre-screened, the signal-to-noise ratio of the database in the subsequent user content preference judgment is improved, the evaluation efficiency and accuracy are optimized, and compared with the noise data screening method based on the content judgment of the conversation preference in the prior art, the noise screening method based on the behavior characteristics can more accurately avoid the inaccurate screening caused by the existence of multiple content themes in a single conversation. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 A flowchart of a data pre-management method based on feature learning.

[0048] Figure 2 A flowchart of a curve fitting step in a data pre-management method based on feature learning.

[0049] Figure 3 A component block diagram of a data pre-management system based on feature learning. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.

[0051] The specific implementation mode of the present application is described in detail below in combination with specific examples.

[0052] As Figure 1 The data pre-management method based on feature learning provided by an embodiment of the present application includes the following steps:

[0053] S10, monitoring the conversation of the user to record the conversation behavior characteristics of the user, the conversation behavior characteristics representing the browsing behavior of the user to the conversation content, including conversation stay records and conversation pointer records;

[0054] S20, performing deviation feature calculation based on the online average stay of each conversation and the conversation stay records to generate the conversation stay characteristics of the current conversation records, the online average stay representing the average browsing stay duration of other users, the deviation feature calculation representing the process of calculating the time difference value and the average browsing stay duration proportion, and the conversation stay characteristics having directionality;

[0055] S30, sequentially statistics several session stay features to make curve fitting, obtain a habit feature curve of the user, and judge a rising turning node of the habit feature curve to mark as a preference change node of the user;

[0056] S40, establish a data pre-screening model based on the preference change node, evaluate session behavior features according to the data pre-screening model to judge whether it is noise data, and the data pre-screening model includes a deviation feature calculation model and a comparison and judgment program based on the preference change node.

[0057] In this embodiment, a data pre-management method based on feature learning is given. By monitoring and recording the content browsing session of the user, the individual behavior characteristics are judged to pre-screen a large number of session data, and the noise data useless for user content preference judgment is pre-screened to improve the signal-to-noise ratio of the database in the subsequent user content preference judgment, optimize the evaluation efficiency and accuracy, and compared with the existing noise data screening method based on content judgment session preference, the noise screening method based on behavior characteristics can more accurately avoid the inaccurate screening caused by multiple content themes in a single session. Data preprocessing is a common data screening and normalization step in the field of big data processing, which aims to make the data participating in evaluation and modeling have a higher signal-to-noise ratio, and at the same time make the format standard and avoid the result error caused by non-standard data. Especially in the current situation of rapid expansion of big data, user preference judgment is widely used to improve the acceptance of big data push in many scenarios. This embodiment gives a method of data pre-screening based on user behavior characteristics to improve the accuracy of user preference judgment. Specifically, this process is completed when the user completes the session interaction. Unlike the existing technology, it also needs to involve the screening process of content theme judgment. First, the user session interaction process is logged to obtain the time length of the user completing the current session and the mouse pointer motion path record during the session process. The motion path record is corresponding to the time information. For the same user, the browsing habit is relatively fixed during the session (there may be some differences due to the difference in the type of data browsed in the session content). That is, for the session content that the user does not like (i.e. noise data for preference judgment), the time process of entering and exiting may be different depending on the time length of determining that he does not like it. But for the session content that the user likes, the time needed to complete the browsing of the session content must depend on the user's browsing speed. Therefore, by calculating the proportion of the difference between the total average time length of a large number of people in the session and the total time length, the noise can be screened. When the data volume is large, the number of sessions whose complete reading time is proportional to the reading speed of the regular interesting content also increases, so an upward turning point can be obtained when the session stay feature is curve-fitted, that is, an approximately horizontal curve segment (excluding the minimum value part, which is the part of directly exiting the session without interest) (at this time, the continuous upward curve caused by the session object with high interest value and repeated reading can be avoided). Through this upward turning point, a screening threshold can be established based on the user's browsing behavior habit, so that the user can make a judgment whether it is noise data at the first time when a new session is executed, and effectively avoid the multi-theme interference based on content screening.

[0058] As Figure 2As shown, as another preferred embodiment of the present application, the step of sequentially counting a plurality of session stay features to perform curve fitting, obtaining a habit feature curve of the user, and judging the rising turning point of the habit feature curve specifically comprises:

[0059] S31, judging and distinguishing the data type of the session corresponding to the session stay feature to obtain a plurality of feature data sets;

[0060] S32, sequentially counting a plurality of session stay features in each of the feature data sets to establish a scatter plot image, and adjacent scatters are distributed at equal intervals on the horizontal coordinate;

[0061] S33, performing curve fitting on the scatter plot image, calculating the change curvature at each scatter point, judging the curve trend based on the curvature, and obtaining a flat interval whose curve trend is represented as horizontal or a preset low inclination angle;

[0062] S34, setting the left end point of the flat interval as a preference change node, and if there are a plurality of flat intervals, setting a plurality of left end points of the first flat interval as a plurality of preference change nodes corresponding to different preference levels.

[0063] In this embodiment, the process of obtaining the rising turning point through curve fitting is further explained. Because there are habitual differences in browsing different types of data (for example, session interaction of text data and session interaction of image data), when performing statistics and curve fitting, it is necessary to first classify. After classification, a plurality of session stay features in each data set are counted to establish a scatter plot. In the statistics, the session stay features are arranged in order based on their size, and the horizontal coordinates are equally spaced. The vertical coordinates are valued based on the size of the session stay features. In the process of curve fitting, there may be multiple flat intervals, and the first flat interval is also the interval with the lowest session stay feature. This part is definitely the part that exits immediately after entering the session, so it is completely noise. For the subsequent multiple flat intervals, they can be used to represent the increase of the user's interest degree and the corresponding multiple behavior feature intervals.

[0064] As another preferred embodiment of the present application, it further comprises the steps of:

[0065] Obtaining the in-out node coordinates of the corresponding session output window, and performing plane alignment on a plurality of session pointer records based on the in-out node coordinates;

[0066] Screening a plurality of aligned session pointer records to obtain the shortest pointer record between the in-out coordinate nodes and taking it as a reference pointer record;

[0067] Based on the reference pointer record, performing discrete evaluation to obtain the aggregation area of the reference pointer record and marking it as a noise area.

[0068] If the session behavior feature is represented as in the noise region, the current session is marked as noise data.

[0069] In this embodiment, the step of noise screening based on session pointer record is supplemented. Based on the understanding of user behavior habits, the user usually quickly browses and then exits after entering the interaction for the session content that is not liked. Therefore, the mouse pointer path in this case is relatively simple, and the active interval is small (which depends on the user's individual behavior habits, so the user individual needs to be fitted by the shortest pointer record). Therefore, a noise region can be determined by the range of the session pointer record. In this region, the user's behavior feature represents a quick exit from the session interaction, that is, the user's preference for the session is low. The session content is judged as noise data for preference.

[0070] As another preferred embodiment of the present application, it further includes the steps of:

[0071] The noise data result of the noise region judgment is cross-judged by the noise data result judged based on the data pre-screening model. If both results represent noise data, it is finally marked as noise data, otherwise, no noise data marking is performed.

[0072] In this embodiment, both the data screening model and the noise region in the foregoing content can be used to judge the session belonging to noise data. However, if they are used alone, there will still be some special cases. For example, when the user thinks about a certain session content, the session pointer record may be stationary for a long time, so that the session pointer record is in the noise region. At this time, if the noise region is used for division, the session will be incorrectly judged as noise data. However, in fact, the user's browsing time of the session is relatively long. Therefore, the judgment result of the data pre-screening model can be cross-verified, so that most of the possible incorrect judgments under the two noise screening methods can be avoided.

[0073] As another preferred embodiment of the present application, it further includes the steps of:

[0074] If the session pointer record exceeds the noise region, the concentration of the session pointer is evaluated, a plurality of content concentration regions are obtained, the content concentration region is used to represent the region with a high proportion of pointer residence time in the session pointer record, and the high residence proportion is selected by region proportion sorting;

[0075] The content concentration region is marked and mapped in the session output window to mark the session content.

[0076] In this embodiment, the concentration evaluation step is supplemented, and based on the conversation pointer record, the user's browsing behavior characteristics can be further judged. For example, when the user performs text conversation interaction, the user may use the mouse pointer to approach or even select the text content of interest or thinking part. Therefore, by evaluating the concentration based on the conversation pointer, multiple concentration areas in the conversation pointer record are obtained, and the concentration areas are mapped with the conversation output window. The conversation content that the user may have high interest in can be marked as concentration before the preference judgment, so that the preference content can be quickly located during the preference judgment.

[0077] As shown in Figure 3 The application also provides a data pre-management system based on feature learning, which comprises:

[0078] A conversation record module 100 is configured to monitor the conversation of the user to record the conversation behavior characteristics of the user, wherein the conversation behavior characteristics represent the browsing behavior of the user on the conversation content, and include conversation stay records and conversation pointer records;

[0079] A feature calculation module 200 is configured to calculate deviation features based on the online average stay and the conversation stay records of each conversation to generate conversation stay features of the current conversation record, wherein the online average stay is used to represent the average browsing stay duration of other users, the deviation feature calculation is used to represent the process of calculating the time difference value and the average browsing stay duration proportion, and the conversation stay features have directionality.

[0080] A fitting and splitting module 300 is configured to sequentially calculate a plurality of conversation stay features to perform curve fitting, obtain a habit feature curve of the user, and judge a rising turning node of the habit feature curve to mark as a preference change node of the user.

[0081] A modeling and screening module 400 is configured to establish a data pre-screening model based on the preference change node, evaluate the conversation behavior characteristics according to the data pre-screening model to judge whether it is noise data, and wherein the data pre-screening model includes a deviation feature calculation model and a comparison and judgment program based on the preference change node.

[0082] As another preferred embodiment of the application, the fitting and splitting module 300 comprises:

[0083] A content classification unit is configured to judge and distinguish the data types of the conversations corresponding to the conversation stay features to obtain a plurality of feature data sets.

[0084] A statistical mapping unit is configured to sequentially calculate a plurality of conversation stay features in each feature data set to establish a scatter plot, and adjacent scatters are distributed at equal intervals on the horizontal coordinate.

[0085] a curve fitting unit, configured to perform curve fitting on the scatter plot image, and calculate a change curvature at each scatter point, determine a curve trend based on the curvature, and obtain a flat interval with a horizontal or preset low inclination angle as a representation of the curve trend;

[0086] a node setting unit, configured to set a left end point of the flat interval as a preference change node, and if there are a plurality of flat intervals, set a plurality of left end points of the first flat interval as a plurality of preference change nodes corresponding to different preference levels.

[0087] As another preferred embodiment of the present application, it further comprises a pointer pre-selection module, specifically comprising:

[0088] a record alignment unit, configured to obtain in-out node coordinates corresponding to a session output window, and perform planar alignment on a plurality of session pointer records based on the in-out node coordinates;

[0089] a reference selection unit, configured to filter the aligned session pointer records, and obtain a shortest pointer record between the in-out coordinate nodes as a reference pointer record;

[0090] a discrete selection unit, configured to perform discrete evaluation based on the reference pointer record, obtain an aggregation area of the reference pointer record, and mark it as a noise area;

[0091] a noise marking unit, configured to mark the current session as noise data if the session behavior feature represents the noise area.

[0092] As another preferred embodiment of the present application, it further comprises a cross-validation unit;

[0093] The cross-validation unit is configured to cross-validate the noise data result of the noise area judgment based on the noise data result of the data pre-filtering model judgment, and if both results represent noise data, mark it as noise data finally, otherwise, do not mark it as noise data.

[0094] As another preferred embodiment of the present application, it further comprises an attention judgment module, specifically comprising:

[0095] a pointer attention judgment unit, configured to perform attention evaluation on the session pointer if the session pointer record exceeds the noise area, and obtain a plurality of content attention areas, the content attention area being used to represent a high residence time proportion area of the pointer in the session pointer record, and the high residence proportion being selected by area proportion sorting;

[0096] a content attention mapping unit, configured to mark the content attention area, and map it on the session output window to mark the session content.

[0097] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, databases, or other media in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0098] Other embodiments of the present disclosure will be apparent to those skilled in the art with the accomplishment of the present disclosure as reflected in the specification and embodiments. The present application is intended to cover any variations, uses, or adaptive changes of the present disclosure following the general principles of the present disclosure and including common knowledge or conventional technical means in the art not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are indicated by the claims.

[0099] It should be understood that the present disclosure is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A data pre-management method based on feature learning, characterized in that, Comprise: Conducting session monitoring on a user to record session behavior characteristics of the user, the session behavior characteristics representing browsing behavior of the user on session content, including session stay records and session pointer records; Calculating deviation features based on online average stay and the session stay records for each session to generate session stay features of current session records, the online average stay being used to represent average browsing stay duration of other users, the deviation feature calculation being used to represent a process of calculating time difference and average browsing stay duration proportion, the session stay features having directionality; Sequentially calculating a plurality of the session stay features to perform curve fitting to obtain a habit feature curve of the user, and judging an upward turning node of the habit feature curve to mark as a preference change node of the user; Establishing a data pre-screening model based on the preference change node, and evaluating session behavior characteristics according to the data pre-screening model to judge whether it is noise data, the data pre-screening model including a deviation feature calculation model and a comparison and judgment program based on the preference change node; The step of sequentially calculating a plurality of the session stay features to perform curve fitting to obtain a habit feature curve of the user, and judging an upward turning node of the habit feature curve specifically comprises: Judging and distinguishing data types of sessions corresponding to the session stay features to obtain a plurality of feature data sets; Sequentially calculating a plurality of session stay features in each of the feature data sets to establish a scatter plot, adjacent scatters being distributed at equal intervals on the horizontal coordinate; Performing curve fitting on the scatter plot and calculating change curvatures at each scatter point, judging curve trends based on the curvatures, and obtaining flat intervals represented as horizontal or preset low inclination angles; Setting the left end point of the flat interval as a preference change node, and if there are a plurality of flat intervals, setting a plurality of left end points of the first flat interval as a plurality of preference change nodes corresponding to different preference levels.

2. The feature learning based data pre-management method of claim 1, wherein, Further comprising steps: Obtaining in-out node coordinates of a session output window, and performing plane alignment on a plurality of session pointer records based on the in-out node coordinates; Screening a plurality of aligned session pointer records to obtain a shortest pointer record between in-out coordinate nodes and take it as a reference pointer record; Performing discrete evaluation based on the reference pointer record to obtain an aggregation area of the reference pointer record and mark it as a noise area; If the session behavior characteristics represent the noise area, the current session is marked as noise data.

3. The feature learning based data pre-management method of claim 2, wherein, Further comprising steps: Cross-judging noise data results of the noise area judging based on noise data results judged by the data pre-screening model, if both results represent noise data, finally marking as noise data, otherwise, not marking as noise data.

4. The data pre-management method based on feature learning according to claim 2, characterized in that, Further comprising steps: If the session pointer record exceeds the noise area, performing concentration evaluation on the session pointer to obtain a plurality of content concentration areas, the content concentration areas being used to represent high stay duration proportion areas of pointers in the session pointer record, and the high stay proportion being selected by area proportion sorting; The content focus area is marked and mapped in the conversation output window to mark the conversation content.

5. A feature learning based data pre-management system, characterized by, Comprise: A conversation recording module for monitoring the conversation of a user to record the conversation behavior characteristics of the user, the conversation behavior characteristics representing the user's browsing behavior of the conversation content, including conversation stay records and conversation pointer records; A feature calculation module for calculating deviation features based on the online average stay and the conversation stay records of each conversation to generate conversation stay features of the current conversation record, the online average stay representing the average browsing stay duration of other users, the deviation feature calculation representing the process of calculating the time difference value and the average browsing stay duration proportion, the conversation stay features having directionality; A fitting and splitting module for sequentially calculating a plurality of conversation stay features to perform curve fitting, obtaining a habit feature curve of the user, and judging the rising turning point of the habit feature curve to mark as a preference change node of the user; A modeling and screening module for establishing a data pre-screening model based on the preference change node, evaluating the conversation behavior characteristics according to the data pre-screening model to determine whether it is noise data, the data pre-screening model including a deviation feature calculation model and a comparison and judgment program based on the preference change node; The fitting and splitting module comprises: A content classification unit for distinguishing the data types of the conversations corresponding to the conversation stay features to obtain a plurality of feature data sets; A statistical mapping unit for sequentially calculating a plurality of conversation stay features in each feature data set to establish a scatter plot, with adjacent scatters distributed at equal intervals on the horizontal coordinate; A curve fitting unit for curve fitting the scatter plot and calculating the change curvature at each scatter point, judging the curve trend based on the curvature, and obtaining a flat interval with a horizontal or pre-set low inclination angle as the curve trend representation; A node setting unit for setting the left end point of the flat interval as a preference change node, and if there are multiple flat intervals, setting multiple left end points of the first flat interval as multiple preference change nodes corresponding to different preference levels.

6. The feature learning based data pre-management system of claim 5, wherein, Further comprising a pointer pre-selection module, specifically comprising: A record alignment unit for obtaining the entry and exit node coordinates corresponding to the conversation output window, and aligning a plurality of conversation pointer records based on the entry and exit node coordinates; A reference selection unit for selecting a plurality of aligned conversation pointer records to obtain the shortest pointer record between the entry and exit coordinate nodes as a reference pointer record; A discrete selection unit for evaluating the reference pointer record based on the reference pointer record to obtain the aggregation area of the reference pointer record and mark it as a noise area; A noise marking unit for marking the current conversation as noise data if the conversation behavior characteristics represent the noise area.

7. The feature learning based data pre-management system of claim 6, wherein, Further comprising a cross-validation unit; The cross-validation unit is used to cross-validate the noise data results of the noise area judgment based on the noise data results of the data pre-screening model judgment, and if both results represent noise data, the final result is marked as noise data, otherwise, no noise data marking is performed.

8. The feature learning based data pre-management system of claim 6, wherein, Further comprising a concentration judgment module, specifically comprising: A pointer concentration judgment unit is used for evaluating the concentration of the conversation pointer if the conversation pointer record exceeds the noise area, obtaining a plurality of content concentration areas, the content concentration area is used to represent the high residence time proportion area of the pointer in the conversation pointer record, and the high residence proportion is selected by area proportion sorting; A content concentration mapping unit is used for marking the content concentration area and mapping in the conversation output window to mark the conversation content.

Citation Information

Patent Citations

  • Information pushing method and system based on smart medical big data

    CN119132630A

  • Personalized content real-time pushing method based on user portrait

    CN120632220A