Feature learning-based data pre-management method and system
By monitoring user session behavior and learning its features, recording session dwell time and pointer records, fitting habit feature curves, identifying preference change nodes, and establishing a data pre-screening model, the problem of inaccurate filtering of noisy data is solved, and the signal-to-noise ratio and accuracy of user preference judgment are improved.
Patent Information
- Application Number
- CN202511383204.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing technologies often fail to accurately filter noisy data, resulting in low efficiency and poor accuracy in judging user preferences, especially when dealing with multi-topic information content.
By monitoring user session behavior, recording session dwell and pointer recording, calculating deviation characteristics, fitting habit feature curves, identifying preference change nodes, establishing a data pre-screening model, and evaluating noisy data.
It improves the signal-to-noise ratio of user preference judgment, optimizes evaluation efficiency and accuracy, accurately avoids inaccurate filtering caused by multiple topics in a single session, and enhances the accuracy of data push.
Smart Images

Figure CN120873532A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of user data screening, specifically a data pre-management method and system based on feature learning. Background Technology
[0002] User preference judgment and data push based on big data are widely used in various fields in the current Internet environment, including but not limited to content platform streaming and user targeting screening on sales platforms. Therefore, the accuracy of user preference judgment is very important for data push. Noise filtering of a large amount of user data before preference judgment can greatly reduce the noise data in preference fitting judgment, making the preference judgment results more accurate.
[0003] Existing noise data filtering and cleaning technologies mainly fall into two categories. First, they use dwell time rules to filter low-quality data; if a user's dwell time on a content page is less than a preset value, the content is considered noise data according to user preferences. Second, they use existing user conversion tags for initial content type screening. However, both methods are affected by the total amount of information and its multi-topic nature, significantly reducing the accuracy of noise data filtering. Ultimately, only a small portion of noise can be removed, while most noise data is categorized and judged during preference fitting. This approach not only increases the efficiency of preference judgment but also lowers its accuracy, thus leaving room for optimization. Summary of the Invention
[0004] The purpose of this invention is to provide a data pre-management method and system based on feature learning to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A data pre-management method based on feature learning, comprising: User session monitoring is performed to record user session behavior characteristics, which characterize the user's browsing behavior of session content, including session dwell records and session pointer records; Based on the average online dwell time of each session and the session dwell time record, a deviation feature calculation is performed to generate the session dwell time feature of the current session record. The average online dwell time is used to characterize the average browsing dwell time of other users. The deviation feature calculation is used to characterize the process of calculating the ratio of the time difference to the average browsing dwell time. The session dwell time feature is directional. The several session dwell characteristics are sequentially statistically analyzed to perform curve fitting, obtain the user's habit feature curve, and determine the rising inflection point of the habit feature curve to mark it as the user's preference change node. A data pre-screening model is established based on the preference change nodes. The conversation behavior characteristics are evaluated according to the data pre-screening model to determine whether the data is noisy. The data pre-screening model includes a deviation feature calculation model and a comparison and judgment procedure based on preference change nodes.
[0006] As a further aspect of the present invention: the step of sequentially statistically analyzing several session dwell characteristics to perform curve fitting, obtaining a user habit feature curve, and determining the rising inflection point of the habit feature curve specifically includes: The data types of the sessions corresponding to the session dwell characteristics are judged and distinguished to obtain multiple feature data sets; Several session dwell features in each of the aforementioned feature datasets are sequentially statistically analyzed to create a scatter plot, with adjacent scatter plots distributed at equal intervals on the horizontal axis. Curve fitting is performed on the scatter image, and the curvature at each scatter point is calculated. The curve trend is determined based on the curvature, and the curve trend is characterized as a horizontal or a flat interval with a preset low tilt angle. The left endpoint of the flat interval is set as a preference change node. If there are multiple flat intervals, the left endpoints of the first flat interval are set as multiple preference change nodes corresponding to different preference levels.
[0007] As a further aspect of the present invention, it also includes the following steps: Obtain the coordinates of the entry and exit nodes of the corresponding session output window, and perform planar alignment of multiple session pointer records based on the coordinates of the entry and exit nodes; Filter the aligned session pointer records to obtain the shortest pointer record between the inbound and outbound coordinate nodes and use it as the base pointer record; Discrete evaluation is performed based on the benchmark pointer records to obtain the clustered regions of the benchmark pointer records, which are then marked as noise regions. If the session behavior characteristics are characterized in the noise region, then the current session is marked as noise data.
[0008] As a further aspect of the present invention, it also includes the following steps: The noise data results judged by the data pre-screening model are cross-judged with the noise data results judged by the noise region. If both results are characterized as noise data, then it is finally marked as noise data; otherwise, it is not marked as noise data.
[0009] As a further aspect of the present invention, it also includes the following steps: If the session pointer record exceeds the noise area, the session pointer is evaluated for focus, and multiple content focus areas are obtained. The content focus areas are used to characterize the areas with a high dwell time ratio of the pointer in the session pointer record. The high dwell time ratio is selected by sorting the area ratio. The content focus area is marked and mapped in the session output window to focus the session content.
[0010] This invention aims to provide a data pre-management system based on feature learning, comprising: The session recording module is used to monitor user sessions and record user session behavior characteristics. These session behavior characteristics characterize the user's browsing behavior on session content, including session dwell time records and session pointer records. The feature calculation module is used to perform deviation feature calculation based on the average online dwell time of each session and the session dwell time record to generate the session dwell time feature of the current session record. The average online dwell time is used to characterize the average browsing dwell time of other users. The deviation feature calculation is used to characterize the process of calculating the ratio of the time difference to the average browsing dwell time. The session dwell time feature is directional. The fitting and splitting module is used to perform sequential statistics on several session dwell features to perform curve fitting, obtain the user's habit feature curve, and determine the rising inflection point of the habit feature curve to mark it as the user's preference change node. The modeling and filtering module is used to establish a data pre-filtering model based on the preference change nodes, and to evaluate the conversation behavior characteristics according to the data pre-filtering model to determine whether it is noisy data. The data pre-filtering model includes a deviation feature calculation model and a comparison and judgment program based on preference change nodes.
[0011] As a further aspect of the present invention: the fitting and splitting module includes: The content classification unit is used to determine and distinguish the data types of the sessions corresponding to the session dwell features in order to obtain multiple feature data sets. The statistical mapping unit is used to perform sequential statistics on several session dwell features in each of the feature data sets to establish a scatter plot, with adjacent scatter plots distributed at equal intervals on the horizontal axis. The curve fitting unit is used to perform curve fitting on the scatter image, calculate the curvature change at each scatter point, determine the curve trend based on the curvature, and obtain a flat interval characterized by the curve trend as horizontal or a preset low angle. The node setting unit is used to set the left endpoint of the flat interval as a preference change node. If there are multiple flat intervals, the left endpoints of the first flat interval are set as multiple preference change nodes corresponding to different preference levels.
[0012] As a further embodiment of the present invention, it also includes a pointer preselection module, specifically comprising: The record alignment unit is used to obtain the coordinates of the entry and exit nodes of the corresponding session output window, and to perform planar alignment of multiple session pointer records based on the coordinates of the entry and exit nodes; The reference selection unit is used to filter several aligned session pointer records, obtain the shortest pointer record between the inbound and outbound coordinate nodes, and use it as the reference pointer record. A discrete selection unit is used to perform discrete evaluation based on a reference pointer record, obtain the clustered region of the reference pointer record, and mark it as a noise region; A noise labeling unit is used to label the current session as noise data if the session behavior characteristics are characterized in the noise region.
[0013] As a further embodiment of the present invention, it also includes a cross-validation unit; The cross-validation unit is used to cross-validate the noise data results of the noise region judgment based on the noise data results judged by the data pre-screening model. If both results are characterized as noise data, then it is finally marked as noise data; otherwise, it is not marked as noise data.
[0014] As a further embodiment of the present invention, it also includes a focus judgment module, specifically comprising: The pointer focus judgment unit is used to evaluate the focus of the session pointer if the session pointer record exceeds the noise area, and obtain multiple content focus areas. The content focus areas are used to characterize the areas with high dwell time of the pointer in the session pointer record. The high dwell time ratio is selected by sorting the area ratio. The content focus mapping unit is used to mark the content focus area and map it in the session output window to mark the session content as focus.
[0015] Compared with the prior art, the beneficial effects of the present invention are: by monitoring and recording users' content browsing sessions, individual behavioral characteristics can be judged to pre-screen a large amount of session data, and noisy data that is useless for judging user content preferences can be pre-screened, thereby improving the signal-to-noise ratio of the database in subsequent user content preference judgment, optimizing evaluation efficiency and accuracy. Moreover, compared with the existing noisy data screening method based on content judgment of session preferences, the noise screening method based on behavioral characteristics can more accurately avoid the situation where the screening is inaccurate due to multiple content topics in a single session. Attached Figure Description
[0016] Figure 1 This is a flowchart of a data pre-management method based on feature learning.
[0017] Figure 2 This is a flowchart of the curve fitting step in a feature-based data pre-management method.
[0018] Figure 3 This is a block diagram of a data pre-management system based on feature learning. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0020] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.
[0021] like Figure 1 The aforementioned data pre-management method based on feature learning, provided in one embodiment of the present invention, includes the following steps: S10, Monitor user sessions to record user session behavior characteristics, which characterize the user's browsing behavior on session content, including session dwell time records and session pointer records; S20, based on the average online dwell time of each session and the session dwell time record, a deviation feature calculation is performed to generate the session dwell time feature of the current session record. The average online dwell time is used to characterize the average browsing dwell time of other users. The deviation feature calculation is used to characterize the process of calculating the ratio of the time difference to the average browsing dwell time. The session dwell time feature has directionality. S30, perform sequential statistics on several of the session dwell features to perform curve fitting, obtain the user's habit feature curve, and determine the rising inflection point of the habit feature curve to mark it as the user's preference change point; S40, establish a data pre-screening model based on the preference change nodes, evaluate the conversation behavior characteristics according to the data pre-screening model to determine whether it is noisy data, the data pre-screening model includes a deviation feature calculation model and a comparison and judgment procedure based on preference change nodes.
[0022] This embodiment presents a data pre-management method based on feature learning. By monitoring and recording users' content browsing sessions, individual behavioral characteristics are determined to pre-screen a large amount of session data. This pre-screening removes noisy data that is useless for judging user content preferences, thereby improving the signal-to-noise ratio of the database in subsequent user content preference judgments and optimizing evaluation efficiency and accuracy. Compared to existing noisy data filtering methods based on content-based session preference judgments, the behavioral feature-based noise filtering method can more accurately avoid inaccurate filtering due to multiple content topics within a single session. Data preprocessing is a common data filtering and normalization step in the field of big data processing technology, aiming to make the data more accurate. Compared to evaluation and modeling data, this method offers a higher signal-to-noise ratio and ensures standardized formatting, avoiding errors caused by non-standard data. Especially in today's rapidly expanding big data landscape, judging user preferences to improve the acceptance of big data pushes is widely used in numerous scenarios. This embodiment presents a method for data pre-screening based on user behavioral characteristics to improve the accuracy of user preference judgment. Specifically, this process is completed when the user completes a conversation interaction. Unlike existing technologies that require a content topic judgment screening process, this method first logs the user's conversation interaction process, obtaining the duration of the current conversation and the mouse pointer movement during the conversation. Motion path records, and these records correspond to time information, indicate that for the same user, their browsing habits during a session are relatively fixed (though they may vary slightly depending on the type of data viewed within the session). That is, for disliked session content (i.e., noise data related to preference judgment), the time taken to enter and exit may vary depending on the duration of determining dislike. However, for liked session content, the time required to complete browsing depends on the user's browsing speed. Therefore, by calculating the ratio of the difference between the average duration of the session and the average duration of a large number of online users to the total duration, noise filtering can be performed. When the data volume is large, this is often... The proportion of conversations where the completion time for content of interest is proportional to the reading speed also increases. Therefore, when fitting the curve of conversation dwell characteristics, an upward turning point can be obtained, which is an approximately horizontal curve segment with a gradual increase (excluding the lowest value part, which is the part where no interest is expressed and the conversation is exited directly). (At this time, the continuous rise of the curve caused by a small number of conversation objects with high interest values are avoided). Through this upward turning point, a filtering threshold can be established based on the user's browsing behavior habits. Thus, when the user finishes a new conversation, the classification of whether it is noise data can be determined immediately, and the interference of multiple topics based on content filtering can be effectively avoided.
[0023] like Figure 2As shown, in another preferred embodiment of the present invention, the step of sequentially statistically analyzing several session dwell features to perform curve fitting, obtaining a user habit feature curve, and determining the rising inflection point of the habit feature curve specifically includes: S31, determine and distinguish the data type of the session corresponding to the session dwell feature to obtain multiple feature data sets; S32, sequentially count several session dwell features in each of the feature data sets to establish a scatter plot, with adjacent scatter plots distributed at equal intervals on the horizontal axis; S33, perform curve fitting on the scatter image and calculate the curvature at each scatter point. Based on the curvature, determine the curve trend and obtain a flat interval characterized by the curve trend as horizontal or a preset low angle. S34, set the left endpoint of the flat interval as a preference change node. If there are multiple flat intervals, then set the multiple left endpoints of the first flat interval as multiple preference change nodes corresponding to different preference levels.
[0024] In this embodiment, the process of obtaining the rising inflection point through curve fitting is further explained. Because there are habitual differences in browsing different types of data (e.g., conversational interaction of text data and conversational interaction of image data), it is necessary to classify them first when performing statistics and curve fitting. After classification, multiple conversation dwell features within each category of dataset are statistically analyzed to establish a scatter plot. In the statistics, the conversation dwell features are arranged in order of magnitude, with equal intervals on the horizontal axis and the vertical axis assigned values based on the magnitude of the conversation dwell features. During the curve fitting process, there may also be multiple plateau intervals. The first plateau interval is also the interval with the lowest conversation dwell feature. This part must be the part where the conversation is entered and then immediately exited, so it is completely noise. The subsequent multiple flat intervals can be used to represent multiple behavioral feature intervals corresponding to the increase in user interest.
[0025] As another preferred embodiment of the present invention, the method further includes the following steps: Obtain the coordinates of the entry and exit nodes of the corresponding session output window, and perform planar alignment of multiple session pointer records based on the coordinates of the entry and exit nodes; Filter the aligned session pointer records to obtain the shortest pointer record between the inbound and outbound coordinate nodes and use it as the base pointer record; Discrete evaluation is performed based on the benchmark pointer records to obtain the clustered regions of the benchmark pointer records, which are then marked as noise regions. If the session behavior characteristics are characterized in the noise region, then the current session is marked as noise data.
[0026] In this embodiment, a noise filtering step based on session pointer records is added. Based on the understanding of user behavior habits, for disliked session content, users usually enter the interaction, quickly browse it briefly, and then exit. Therefore, the mouse pointer path in this case is relatively simple, and the activity range is small (the specific range depends on the user's individual behavior habits, so it is necessary to fit the individual user through the shortest pointer record). Therefore, a noise area can be defined by the concentration range of session pointer records. Within this area, the user's behavior characteristics indicate that they quickly exit the session interaction, that is, the user has a low liking for the session, and the session content is judged as noise data for preference.
[0027] As another preferred embodiment of the present invention, the method further includes the following steps: The noise data results judged by the data pre-screening model are cross-judged with the noise data results judged by the noise region. If both results are characterized as noise data, then it is finally marked as noise data; otherwise, it is not marked as noise data.
[0028] In this embodiment, both the data filtering model and the noise region mentioned above can be used to determine whether a session is noisy data. However, if the two are used separately, there will still be some special cases. For example, when a user is thinking about a session, the session pointer record may be static for a long time, thus making the session pointer record in the noise region. If the session is divided by the noise region, it will be incorrectly judged as noisy data. However, in reality, the user's browsing time for the session is relatively long. Therefore, the judgment result of the data pre-filtering model can be cross-validated to avoid the erroneous judgment that may occur under the two noise filtering methods in most cases.
[0029] As another preferred embodiment of the present invention, the method further includes the following steps: If the session pointer record exceeds the noise area, the session pointer is evaluated for focus, and multiple content focus areas are obtained. The content focus areas are used to characterize the areas with a high dwell time ratio of the pointer in the session pointer record. The high dwell time ratio is selected by sorting the area ratio. The content focus area is marked and mapped in the session output window to focus the session content.
[0030] In this embodiment, a focus evaluation step is added. Based on the session pointer record, further judgment can be made on the user's browsing behavior characteristics. For example, when the user is engaged in text-based session interaction, they may use the mouse pointer to move close to or even select the text content that they are interested in or thinking about. Therefore, focus evaluation is performed through the session pointer to obtain multiple focus areas in the session pointer record and map the focus areas to the session output window. Before making preference judgment, the session content that the user may be highly interested in can be marked with focus, so that the preferred content can be quickly located during preference judgment.
[0031] like Figure 3 As shown, the present invention also provides a data pre-management system based on feature learning, which includes: The session recording module 100 is used to monitor user sessions and record user session behavior characteristics, which characterize the user's browsing behavior on session content, including session dwell records and session pointer records. The feature calculation module 200 is used to perform deviation feature calculation based on the average online dwell time of each session and the session dwell time record to generate the session dwell time feature of the current session record. The average online dwell time is used to characterize the average browsing dwell time of other users. The deviation feature calculation is used to characterize the process of calculating the ratio of the time difference to the average browsing dwell time. The session dwell time feature has directionality. The fitting and splitting module 300 is used to perform sequential statistics on several session dwell features to perform curve fitting, obtain the user's habit feature curve, and determine the rising inflection point of the habit feature curve to mark it as the user's preference change node. The modeling and filtering module 400 is used to establish a data pre-filtering model based on the preference change nodes, and to evaluate the conversation behavior characteristics according to the data pre-filtering model to determine whether it is noisy data. The data pre-filtering model includes a deviation feature calculation model and a comparison and judgment program based on preference change nodes.
[0032] In another preferred embodiment of the present invention, the fitting and splitting module 300 includes: The content classification unit is used to determine and distinguish the data types of the sessions corresponding to the session dwell features in order to obtain multiple feature data sets. The statistical mapping unit is used to perform sequential statistics on several session dwell features in each of the feature data sets to establish a scatter plot, with adjacent scatter plots distributed at equal intervals on the horizontal axis. The curve fitting unit is used to perform curve fitting on the scatter image, calculate the curvature change at each scatter point, determine the curve trend based on the curvature, and obtain a flat interval characterized by the curve trend as horizontal or a preset low angle. The node setting unit is used to set the left endpoint of the flat interval as a preference change node. If there are multiple flat intervals, the left endpoints of the first flat interval are set as multiple preference change nodes corresponding to different preference levels.
[0033] In another preferred embodiment of the present invention, a pointer preselection module is further included, specifically comprising: The record alignment unit is used to obtain the coordinates of the entry and exit nodes of the corresponding session output window, and to perform planar alignment of multiple session pointer records based on the coordinates of the entry and exit nodes; The reference selection unit is used to filter several aligned session pointer records, obtain the shortest pointer record between the inbound and outbound coordinate nodes, and use it as the reference pointer record. A discrete selection unit is used to perform discrete evaluation based on a reference pointer record, obtain the clustered region of the reference pointer record, and mark it as a noise region; A noise labeling unit is used to label the current session as noise data if the session behavior characteristics are characterized in the noise region.
[0034] As another preferred embodiment of the present invention, a cross-validation unit is also included; The cross-validation unit is used to cross-validate the noise data results of the noise region judgment based on the noise data results judged by the data pre-screening model. If both results are characterized as noise data, then it is finally marked as noise data; otherwise, it is not marked as noise data.
[0035] In another preferred embodiment of the present invention, a focus determination module is also included, specifically comprising: The pointer focus judgment unit is used to evaluate the focus of the session pointer if the session pointer record exceeds the noise area, and obtain multiple content focus areas. The content focus areas are used to characterize the areas with high dwell time of the pointer in the session pointer record. The high dwell time ratio is selected by sorting the area ratio. The content focus mapping unit is used to mark the content focus area and map it in the session output window to mark the session content as focus.
[0036] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0037] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the disclosure in the specification and embodiments. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0038] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A data pre-management method based on feature learning, characterized in that, Include: User session monitoring is performed to record user session behavior characteristics, which characterize the user's browsing behavior of session content, including session dwell records and session pointer records; Based on the average online dwell time of each session and the session dwell time record, a deviation feature calculation is performed to generate the session dwell time feature of the current session record. The average online dwell time is used to characterize the average browsing dwell time of other users. The deviation feature calculation is used to characterize the process of calculating the ratio of the time difference to the average browsing dwell time. The session dwell time feature is directional. The several session dwell characteristics are sequentially statistically analyzed to perform curve fitting, obtain the user's habit feature curve, and determine the rising inflection point of the habit feature curve to mark it as the user's preference change node. A data pre-screening model is established based on the preference change nodes. The conversation behavior characteristics are evaluated according to the data pre-screening model to determine whether the data is noisy. The data pre-screening model includes a deviation feature calculation model and a comparison and judgment procedure based on preference change nodes.
2. The data pre-management method based on feature learning according to claim 1, characterized in that, The step of sequentially statistically analyzing several session dwell features to perform curve fitting, obtaining a user habit feature curve, and determining the rising inflection point of the habit feature curve specifically includes: The data types of the sessions corresponding to the session dwell characteristics are judged and distinguished to obtain multiple feature data sets; Several session dwell features in each of the aforementioned feature datasets are sequentially statistically analyzed to create a scatter plot, with adjacent scatter plots distributed at equal intervals on the horizontal axis. Curve fitting is performed on the scatter image, and the curvature at each scatter point is calculated. The curve trend is determined based on the curvature, and a flat interval characterized by the curve trend as horizontal or a preset low angle is obtained. The left endpoint of the flat interval is set as a preference change node. If there are multiple flat intervals, the left endpoints of the first flat interval are set as multiple preference change nodes corresponding to different preference levels.
3. The data pre-management method based on feature learning according to claim 2, characterized in that, It also includes the following steps: Obtain the coordinates of the entry and exit nodes of the corresponding session output window, and perform planar alignment of multiple session pointer records based on the coordinates of the entry and exit nodes; Filter the aligned session pointer records to obtain the shortest pointer record between the inbound and outbound coordinate nodes and use it as the base pointer record; Discrete evaluation is performed based on the benchmark pointer records to obtain the clustered regions of the benchmark pointer records, which are then marked as noise regions; If the session behavior characteristics are characterized in the noise region, then the current session is marked as noise data.
4. The data pre-management method based on feature learning according to claim 3, characterized in that, It also includes the following steps: The noise data results judged by the data pre-screening model are cross-judged with the noise data results judged by the noise region. If both results are characterized as noise data, then it is finally marked as noise data; otherwise, it is not marked as noise data.
5. The data pre-management method based on feature learning according to claim 3, characterized in that, It also includes the following steps: If the session pointer record exceeds the noise area, the session pointer is evaluated for focus, and multiple content focus areas are obtained. The content focus areas are used to characterize the areas with a high dwell time ratio of the pointer in the session pointer record. The high dwell time ratio is selected by sorting the area ratio. The content focus area is marked and mapped in the session output window to focus the session content.
6. A data pre-management system based on feature learning, characterized in that, Include: The session recording module is used to monitor user sessions and record user session behavior characteristics. These session behavior characteristics characterize the user's browsing behavior on session content, including session dwell time records and session pointer records. The feature calculation module is used to perform deviation feature calculation based on the average online dwell time of each session and the session dwell time record to generate the session dwell time feature of the current session record. The average online dwell time is used to characterize the average browsing dwell time of other users. The deviation feature calculation is used to characterize the process of calculating the ratio of the time difference to the average browsing dwell time. The session dwell time feature is directional. The fitting and splitting module is used to perform sequential statistics on several session dwell features to perform curve fitting, obtain the user's habit feature curve, and determine the rising inflection point of the habit feature curve to mark it as the user's preference change node. The modeling and filtering module is used to establish a data pre-filtering model based on the preference change nodes, and to evaluate the conversation behavior characteristics according to the data pre-filtering model to determine whether it is noisy data. The data pre-filtering model includes a deviation feature calculation model and a comparison and judgment program based on preference change nodes.
7. The data pre-management system based on feature learning according to claim 6, characterized in that, The fitting and splitting module includes: The content classification unit is used to determine and distinguish the data types of the sessions corresponding to the session dwell features in order to obtain multiple feature data sets. The statistical mapping unit is used to perform sequential statistics on several session dwell features in each of the feature data sets to establish a scatter plot, with adjacent scatter plots distributed at equal intervals on the horizontal axis. The curve fitting unit is used to perform curve fitting on the scatter image, calculate the curvature change at each scatter point, determine the curve trend based on the curvature, and obtain a flat interval characterized by the curve trend as horizontal or a preset low tilt angle. The node setting unit is used to set the left endpoint of the flat interval as a preference change node. If there are multiple flat intervals, the left endpoints of the first flat interval are set as multiple preference change nodes corresponding to different preference levels.
8. The data pre-management system based on feature learning according to claim 7, characterized in that, It also includes a pointer preselection module, specifically including: The record alignment unit is used to obtain the coordinates of the entry and exit nodes of the corresponding session output window, and to perform planar alignment of multiple session pointer records based on the coordinates of the entry and exit nodes; The reference selection unit is used to filter several aligned session pointer records, obtain the shortest pointer record between the inbound and outbound coordinate nodes, and use it as the reference pointer record. A discrete selection unit is used to perform discrete evaluation based on a reference pointer record, obtain the clustered region of the reference pointer record, and mark it as a noise region; A noise labeling unit is used to label the current session as noise data if the session behavior characteristics are characterized in the noise region.
9. The data pre-management system based on feature learning according to claim 8, characterized in that, It also includes cross-validation units; The cross-validation unit is used to cross-validate the noise data results of the noise region judgment based on the noise data results judged by the data pre-screening model. If both results are characterized as noise data, then it is finally marked as noise data; otherwise, it is not marked as noise data.
10. The data pre-management system based on feature learning according to claim 8, characterized in that, It also includes a focus judgment module, specifically including: The pointer focus judgment unit is used to evaluate the focus of the session pointer if the session pointer record exceeds the noise area, and obtain multiple content focus areas. The content focus areas are used to characterize the areas with high dwell time of the pointer in the session pointer record. The high dwell time ratio is selected by sorting the area ratio. The content focus mapping unit is used to mark the content focus area and map it in the session output window to mark the session content as focus.
Citation Information
Patent Citations
Two-channel graph neural network session recommendation method and system based on time interval
CN117828181A
Information pushing method and system based on smart medical big data
CN119132630A
Personalized content real-time pushing method based on user portrait
CN120632220A
Method and system for extracting user behavior features to personalize recommendations
US20150088911A1