Point of interest extraction method and apparatus, terminal device, and storage medium

CN118245662BActive Publication Date: 2026-07-21GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2022-12-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Without the user's explicit notification, existing technologies struggle to accurately perceive and explicitly represent a user's various types of interests over a given period, especially in short-term switching scenarios between super apps, and existing methods may infringe on user privacy.

Method used

By acquiring target usage data from terminal devices, including APP usage data, description data, and screen-on and screen-off information, vectorization is performed, and similarity calculations are performed with historical data to determine target interest points. Pre-trained models such as BERT or word2vec are used for text vectorization, and combined with text clustering and similarity algorithms, user interest points are extracted.

Benefits of technology

It can accurately perceive and represent user interests without requiring users to actively inform it, thereby improving user experience, providing rich descriptions of interests for downstream businesses, and avoiding privacy violations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118245662B_ABST
    Figure CN118245662B_ABST
Patent Text Reader

Abstract

The application provides a point of interest extraction method, device, terminal equipment and storage medium. In the technical solution of the application, target use data of a terminal equipment in a target time period is acquired, the target use data including use data of at least one APP, description data of the at least one APP and state data including screen-on information and screen-off information of the terminal equipment; then the target use data is subjected to vectorization processing, and the vectorized target use data is subjected to similarity calculation with a plurality of preset vectorized data to determine target vectorized data; finally, a preset point of interest corresponding to the target vectorized data is determined as a target point of interest in the target time period. The point of interest extraction method of the application can greatly facilitate accurate perception of a user's point of interest in a certain time period without active user notification, and provide rich user point of interest description for downstream business parties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to methods, apparatus, terminal devices and storage media for extracting points of interest. Background Technology

[0002] With the development of communication technology, terminals include more and more applications (APPs). Terminals can extract users' points of interest over a certain period of time based on their APP usage records, providing convenience for downstream business partners in the terminal.

[0003] However, most current apps are super apps, which can be considered as a collection of multiple functions. For example, WeChat includes chat, games, short videos, food delivery, and ride-hailing, while Meituan includes food delivery, reviews, ride-hailing, and movie tickets. Users' short-term interests often involve switching between apps with similar functions. For instance, when a user becomes interested in hailing a ride, they will switch between WeChat, Meituan, and Didi Chuxing in a short period of time.

[0004] Therefore, how to accurately perceive and explicitly represent a user's interests over a certain period of time without the user actively informing them has become a technical problem that this application urgently needs to solve. Summary of the Invention

[0005] This application provides a method, apparatus, terminal device, and storage medium for extracting points of interest, which can accurately perceive and explicitly represent a user's points of interest within a certain period of time without the user's active notification, providing rich descriptions of user points of interest for downstream business parties.

[0006] Firstly, this application provides a method for extracting points of interest (POIs). The method includes: acquiring target usage data of a terminal device during a target time period, the target usage data including usage data of at least one application (APP) within the target time period, description data of the at least one APP, and status data including screen-on and screen-off information of the terminal device; vectorizing the target usage data and calculating the similarity between the vectorized target usage data and multiple preset vectorized data to determine target vectorized data, wherein the target vectorized data is the vectorized data with the highest similarity to the vectorized target usage data among the multiple preset vectorized data, and the multiple preset vectorized data are obtained by vectorizing historical usage data prior to the target time period; and determining the preset POIs corresponding to the target vectorized data as the target POIs for the target time period, wherein the multiple preset vectorized data correspond to multiple preset POIs respectively.

[0007] In the above technical solution, the terminal device calculates the similarity between the vectorized target usage data and multiple preset vectorized data, and then finally determines the target interest points for the target time period. Compared with the existing technology, it can accurately perceive the user's interest points within a certain period of time without the user actively informing the user, making it convenient to extract the user's interest points that change over time or periodically, while providing rich descriptions of user interest points for downstream business parties and improving user experience.

[0008] Secondly, this application provides an interest point extraction device, the device comprising: an acquisition module, configured to acquire target usage data of a terminal device during a target time period, the target usage data including usage data of at least one application (APP) during the target time period, description data of the at least one APP, and status data including screen-on information and screen-off information of the terminal device; a processing module, configured to perform vectorization processing on the target usage data, and perform similarity calculation between the vectorized target usage data and multiple preset vectorized data to determine target vectorized data, the target vectorized data being the vectorized data with the highest similarity among the multiple preset vectorized data, the multiple preset vectorized data being obtained by vectorizing historical usage data prior to the target time period; and a determination module, configured to determine the preset interest point corresponding to the target vectorized data as the target interest point of the target time period, the multiple preset vectorized data corresponding to multiple preset interest points respectively.

[0009] Thirdly, this application provides a terminal device, including a processor and a memory, wherein the memory is used to store code instructions; and the processor is used to execute the code instructions to implement the method in the first aspect above.

[0010] Optionally, there may be one or more processors and one or more memories.

[0011] Alternatively, the memory can be integrated with the processor, or the memory can be set up separately from the processor.

[0012] In specific implementation, the memory can be a non-transitory memory, such as read-only memory (ROM), which can be integrated with the processor on the same chip or set on different chips. The embodiments of this application do not limit the type of memory or the way the memory and processor are set.

[0013] Fourthly, this application provides a computer-readable storage medium storing a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the methods described in the first aspect above.

[0014] Fifthly, this application provides a computer program product comprising: a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the method described in the first aspect above. Attached Figure Description

[0015] Figure 1 A schematic diagram of the system architecture of a terminal device provided in one embodiment of this application;

[0016] Figure 2 A flowchart illustrating a method for extracting points of interest according to an embodiment of this application;

[0017] Figure 3 A flowchart illustrating a method for extracting points of interest according to another embodiment of this application;

[0018] Figure 4 A flowchart illustrating a method for extracting points of interest according to yet another embodiment of this application;

[0019] Figure 5 A structural schematic diagram of an interest point extraction device provided in one embodiment of this application;

[0020] Figure 6 This is a structural schematic diagram of an apparatus provided for another embodiment of this application. Detailed Implementation

[0021] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0022] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. For example, "first instruction" and "second instruction" are used to distinguish different user instructions and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.

[0023] It should be noted that, in this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0024] Furthermore, "at least one" refers to one or more, while "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0025] With the development of communication technology, terminals have integrated more and more functions, resulting in a growing number of corresponding applications (APPs) in the terminal's system function list. The terminal can extract the user's points of interest over a certain period of time based on the user's APP usage records, providing convenience for downstream business parties in the terminal.

[0026] However, most current apps are super apps, which can be considered as a collection of multiple functions. For example, WeChat includes chat, games, short videos, food delivery, and ride-hailing, while Meituan includes food delivery, reviews, ride-hailing, and movie tickets. Users' short-term interests often involve switching between apps with similar functions. For instance, when a user becomes interested in hailing a ride, they will switch between WeChat, Meituan, and Didi Chuxing within a short period. Therefore, effectively perceiving and explicitly expressing a user's interests over a specific time period without the user actively informing them is a significant challenge.

[0027] In current related technologies, terminal devices can use classification algorithms to merge search records that are determined to have the same intent or similar semantics into the same intent block based on the user's search history in the browser, and then output each intent block, which is the user's set of interests.

[0028] However, since a super app typically includes multiple different types of user interests, and these interests vary significantly, the aforementioned approach, which assumes a search record contains only one interest and then merges multiple search records into a single intent block to extract user intent, provides coarse-grained intent extraction. This approach is unsuitable for scenarios where super apps contain diverse user interests, and the extracted results cannot be represented semantically. Furthermore, with users increasingly prioritizing privacy in recent years, the data used in the aforementioned approach—which relies on users directly searching for text—is likely to alienate users.

[0029] Therefore, how to accurately perceive and explicitly represent a user's interests within a certain time period without the user actively informing them has become a technical problem that this application urgently needs to solve.

[0030] In view of this, embodiments of this application provide a method, apparatus, terminal device, and storage medium for extracting points of interest. In the technical solution of this application, the terminal device acquires target usage data of the terminal device within a target time period. The target usage data includes usage data of at least one application (APP) within the target time period, description data of at least one APP, and status data including screen-on and screen-off information of the terminal device. The target usage data is vectorized, and the similarity of the vectorized target usage data with multiple preset vectorized data is calculated to determine target vectorized data. The target vectorized data is the vectorized data with the highest similarity to the vectorized target usage data among the multiple preset vectorized data. The multiple preset vectorized data are obtained by vectorizing historical usage data prior to the target time period. Preset points of interest corresponding to the target vectorized data are determined as target points of interest for the target time period, and the multiple preset vectorized data correspond to multiple preset points of interest. The point of interest extraction method of this application can greatly facilitate the accurate perception of a user's points of interest within a certain time period without the user actively informing them, providing rich descriptions of user points of interest for downstream businesses.

[0031] It should be understood that the terminal devices involved in the embodiments of this application may be mobile phones, tablet computers, laptops, handheld computers, mobile internet devices (MID), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, smart screens, artificial intelligence (AI) speakers, headphones, terminals in industrial control, terminals in self-driving, terminals in remote medical surgery, terminals in smart grids, terminals in transportation safety, terminals in smart cities, terminals in smart homes, personal digital assistants (PDAs), etc., and the embodiments of this application are not limited to these.

[0032] For example, Figure 1 This is a schematic diagram of the system architecture of the terminal device provided in an embodiment of this application. Figure 1 As shown, the terminal device includes components such as a processor 110, a memory 120, a transceiver 130, a display unit 140, an input unit 150, a sensor 160, an audio circuit 170, and a power module 180.

[0033] The processor 110 is the control center of the terminal device. It connects various parts of the terminal device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 120, and by calling data stored in the memory 120, it performs various functions of the terminal device and processes data, thereby providing overall monitoring of the terminal device. Optionally, the processor 110 may include one or more processing units; optionally, the processor 110 may integrate an application processor, which mainly handles operating devices, user interfaces, and applications, etc. Of course, it may also include other processors, which are not listed here.

[0034] The memory 120 can be used to store software programs and modules. The processor 110 executes various functional applications and data processing of the terminal device by running the software programs and modules stored in the memory 120. The memory 120 mainly includes a program storage area and a data storage area. The program storage area can store the operating device and application programs required for at least one function (such as sound playback function, image playback function, etc.). The data storage area can store data created according to the use of the terminal device (such as audio data, phone book, etc.). In addition, the memory 120 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0035] Transceiver 130 can provide solutions for wireless communication applications on terminal devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. Transceiver 130 can be one or more devices integrating at least one communication processing module; for example, it can integrate an antenna with a baseband processor, or it can integrate an antenna with a modem processor, etc., without limitation.

[0036] The display unit 140 can be used to display information input by the user or information provided to the user, as well as various menus of the terminal device. The display unit 140 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like, and is not limited thereto.

[0037] The input unit 150 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the terminal device. Specifically, the input unit 150 can collect user operations on or near it and drive corresponding connection devices according to a pre-set program. Furthermore, the input unit 150 may include a touch panel, which can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave touch panels. In addition to the touch panel, the input unit 150 may also include other input devices. Specifically, other input devices may include, but are not limited to, one or more of the following: function keys (such as volume control buttons, power buttons, etc.), trackballs, joysticks, etc.

[0038] The terminal device may also include at least one sensor 160, such as a gyroscope sensor, a motion sensor, and other sensors. The motion sensor may include an accelerometer sensor to detect the magnitude of acceleration in various directions. When stationary, it can detect the magnitude and direction of gravity and can be used for applications that identify the terminal device's posture, such as landscape / portrait switching, related games, and magnetometer posture calibration. Other sensors that the terminal device may also be configured with, such as pressure gauges, barometers, hygrometers, thermometers, infrared sensors, and fingerprint sensors, will not be described in detail here.

[0039] The audio circuit 170 may include a speaker and a microphone, providing an audio interface between the user and the terminal device. The audio circuit 170 can convert received audio data into electrical signals and transmit them to the speaker, where the speaker converts them into sound signals for output. On the other hand, the microphone converts collected sound signals into electrical signals, which are received by the audio circuit 170, converted into audio data, and then processed by the processor 110 before being sent to, for example, another terminal device via the video circuit, or the audio data can be output to the memory 120 for further processing.

[0040] The terminal device also includes a power module 180 that supplies power to the various components. Optionally, the power module 180 can be logically connected to the processor 110 through a power management device, thereby enabling the power management device to manage functions such as charging, discharging, and power consumption.

[0041] Although not shown, the terminal device may also include a camera. Optionally, the camera may be positioned in the front or rear of the terminal device, and this application embodiment does not limit this.

[0042] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device. In other embodiments of this application, the terminal device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0043] To make the objectives and technical solutions of this application clearer and more intuitive, the method, apparatus, terminal device, and storage medium for extracting points of interest provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0044] Please refer to Figure 2 This is a flowchart illustrating a method for extracting points of interest according to an embodiment of this application. This method can be applied to, for example... Figure 1 The terminal device shown below will be described in detail. Figure 2 The flowchart illustrates the various steps in the method shown, including:

[0045] S201, Obtain target usage data of the terminal device during the target time period. The target usage data includes usage data of at least one APP during the target time period, description data of at least one APP, and status data including screen-on information and screen-off information of the terminal device.

[0046] It should be understood that app usage data primarily refers to the user's operation records when opening or closing apps on the terminal device. For example, a user opens the Douyin app on their terminal device, browses for 3 minutes, closes the Douyin app, and then opens the Kuaishou app to continue browsing, and so on. App description data refers to the app's relevant introductory data, which is pre-stored in the terminal device's database. This typically includes data such as the app's description and reviews from the app store on the terminal device. Screen-on and screen-off status data refers to the operation records of the terminal device's screen being turned on or off. For example, a user turns on the terminal device's screen, performs some operations on the device, and then turns it off.

[0047] S202, the target usage data is vectorized, and the similarity between the vectorized target usage data and multiple preset vectorized data is calculated to determine the target vectorized data. The target vectorized data is the vectorized data with the highest similarity to the vectorized target usage data among the multiple preset vectorized data. The multiple preset vectorized data are obtained by vectorizing historical usage data before the target time period.

[0048] It should be understood that vectorization can employ pre-trained models (bidirectional encoder representations from transformers, BERT), word2vec, and other text vectorization techniques. The core idea of ​​word2vec is to use a neural network to train a word's context to obtain its vectorized representation. Training methods include predicting the center word from nearby words (CBOW) and predicting nearby words from the center word (Skip-gram). BERT essentially learns a good feature representation for words by running a self-supervised learning method on massive amounts of corpus data. Self-supervised learning refers to supervised learning performed on unlabeled data. In specific natural language processing (NLP) tasks, we can directly use BERT's feature representations as word embedding features for that task. Therefore, BERT provides a model for transfer learning in other tasks; this model can be fine-tuned or fixed as a feature extractor depending on the task. This application does not limit the specific vectorization methods used.

[0049] It should also be understood that multiple preset vectorized data can be obtained through preprocessing or updated in real time, and this application does not limit this.

[0050] S203, determine the preset interest points corresponding to the target vectorized data as the target interest points for the target time period, and the multiple preset vectorized data correspond to the multiple preset interest points respectively.

[0051] It should be understood that each preset vectorized data corresponds to a preset point of interest, and the preset point of interest corresponding to the target vectorized data determined in step S202 is determined as the target point of interest for the target time period.

[0052] In this embodiment, the terminal device calculates the similarity between the vectorized target usage data and multiple preset vectorized data, thereby ultimately determining the target interest points for the target time period. Compared with the prior art, this greatly facilitates the accurate perception of a user's interest points within a certain period of time without the user actively informing the user. It also facilitates the extraction of user interest points that change over time or periodically, while providing rich descriptions of user interest points for downstream businesses, thus improving the user experience.

[0053] Based on the above embodiments, Figure 3 A flowchart of a method for extracting points of interest provided in another embodiment of this application is shown. Figure 3 In the illustrated embodiment, taking as an example how a terminal device determines multiple preset vectorized data and multiple preset interest points corresponding to the multiple preset vectorized data, the following is a detailed explanation. Figure 3The flowchart illustrates the various steps in the method shown, including:

[0054] S301, Obtain historical usage data, which includes usage data of at least one APP within at least one historical time period, description data of at least one APP, and status data including screen-on information and screen-off information of the terminal device.

[0055] It should be understood that before acquiring historical usage data, the terminal device can first initialize itself to eliminate the influence of previous data on the collection of historical usage data and improve the accuracy of the collected data samples.

[0056] For example, taking a historical time period from 7:30 to 8:30, and at least one app being WeChat and Honor of Kings, assume that the app usage data includes: ID_timestamp, APP(0, 2022-10-24 08:00, WeChat; 1, 2022-10-24 08:02, Honor of Kings); the terminal device status data includes: ID_timestamp, screen status screen(0, 2022-10-24 07:59, screen on; 1, 2022-10-24 08:30, screen off); the app description data includes: application description information app desc(WeChat, description information; Honor of Kings, description information).

[0057] It should be noted that since the description information for applications is generally quite long, including an introduction to the application, reviews, etc., detailed information is not given in the examples above; instead, it is summarized using the description information.

[0058] S302 breaks down historical usage data into multiple sub-data.

[0059] It should be understood that the terminal device divides the historical usage data into multiple sessions based on the status data in the historical usage data. Each session corresponds to a screen-on event or a screen-off event. Each session includes apps that are opened sequentially in the same screen-on event with an adjacent time interval of less than a preset time interval threshold.

[0060] A session refers to the time interval during which an end user communicates with the interactive system. It typically refers to the time elapsed from registering to logging out of the system, and may also include some operational space if needed. For example, if APP1 and APP2 are operated on during the same screen-on event, and the time difference between their operations is less than a preset time interval threshold (e.g., 10 seconds), then APP1 and APP2 are considered to be in the same session; otherwise, two separate sessions are created.

[0061] Furthermore, sub-data corresponding to each session is obtained. The sub-data includes usage data of at least one APP included in the corresponding session and description data of at least one APP, resulting in multiple sets of sub-data.

[0062] In other words, each session includes at least one app, and each piece of sub-data includes one session.

[0063] For example, taking a historical time period from 7:30 to 8:30, and at least one app being WeChat and Honor of Kings, then Honor of Kings and WeChat belong to the same session.

[0064] S303 converts multiple sub-data into vector representations to obtain multiple preset vectorized data.

[0065] In one possible implementation, based on at least one app in each session and description data of at least one app, a session description text corresponding to each session is determined, thereby obtaining multiple session description texts.

[0066] It should be understood that the process involves extracting a list of all apps in each session, merging all app names and app description texts in that list into a single text string, and obtaining the corresponding session description text.

[0067] Furthermore, using text vectorization technology, multiple session description texts are converted into vector representations to obtain multiple preset vectorized data.

[0068] For example, taking a historical time period from 7:30 to 8:30, and at least one app being WeChat and Honor of Kings, the conversation description text corresponding to the conversation including WeChat and Honor of Kings would be application name_description information apptitledesc(WeChat, description information; Honor of Kings, description information). Using text vectorization technology, the conversation description text apptitle desc(WeChat, description information; Honor of Kings, description information) is converted into a vector representation. For example, the commonly used word vectorization technology word2vec converts each word into a vector. For example, the vector value corresponding to "I" might be [0.3, 0.5, 0.7, 0.9, -0.2, 0.03], ultimately resulting in a description text vector representation that includes many word vectors.

[0069] S304, According to the text clustering algorithm, merge multiple preset vectorized data to obtain at least one preset vectorized data cluster.

[0070] In essence, text clustering algorithms compare the similarity of several texts and group those with high similarity into one category. Currently, text clustering algorithms are mainly classified into the following categories: partition-based clustering algorithms, hierarchical clustering algorithms, density-based clustering algorithms, grid-based clustering algorithms, model-based clustering algorithms, and fuzzy clustering algorithms.

[0071] Among text clustering algorithms, partitioning-based algorithms are the simplest, including K-means, Sing-Pass incremental clustering, and Partitioning around Mediats (PAM). They are highly applicable to most data types, computationally simple and efficient, and have low space complexity. Hierarchical clustering (HC), also known as tree clustering, primarily merges or splits a sample set into more cohesive or finer-grained sub-sample sets, ultimately forming a hierarchical tree. Unlike K-means, hierarchical clustering does not require a pre-defined number of clusters; it only requires iterative iteration to reach the clustering condition or the required number of iterations. A classic example of hierarchical partitioning-based clustering is the Chameleon algorithm. Density-based clustering algorithms first identify high-density points, then connect nearby high-density points to form clusters. These algorithms are robust and applicable to clusters of arbitrary shapes; however, their accuracy is highly dependent on parameter settings, limiting their practicality. Compared to other clustering algorithms, grid-based clustering algorithms start from space rather than a plane. In this space, a finite number of grids represent data, and clustering is the process of merging grids according to certain rules. Because grid-based clustering algorithms process data independently, relying only on the number of units in each dimension of the grid structure, they are very fast. However, this algorithm is very sensitive to parameters; the speed comes at the cost of low accuracy, and it usually needs to be used in conjunction with other clustering algorithms. Model-based clustering algorithms assume each cluster is a model and then find the data that best fits that model. There are usually two methods: probability-based and neural network-based. Fuzzy clustering algorithms mainly aim to overcome the either-or classification problem. Their main idea is to use fuzzy set theory as a mathematical foundation and perform cluster analysis using fuzzy mathematics methods. The advantage of this method is that it works well for normally distributed sample data. However, this algorithm relies heavily on initial cluster centers, requiring multiple iterations to find the optimal points, which greatly increases the time complexity for large-scale data samples. This application does not limit the types of text clustering algorithms.

[0072] In this step, the terminal device uses a text clustering algorithm to merge similar preset vectorized data from multiple preset vectorized data to obtain at least one preset vectorized data cluster, and extracts the center coordinates corresponding to each preset vectorized data cluster. The specific text clustering process can refer to the text clustering methods in the prior art, and will not be elaborated here.

[0073] S305, based on a text similarity algorithm and the conversation description text corresponding to at least one preset vectorized data cluster, determine the interest points corresponding to at least one preset vectorized data cluster.

[0074] It should be understood that in text similarity algorithms, text similarity refers to the degree of similarity between two or more entities (words, short texts, documents). It can also be understood as the commonalities or differences between texts; the greater the commonalities and the smaller the differences, the higher the similarity. The similarity is highest when two text segments are completely identical. Common methods include semantic understanding-based text similarity algorithms and word-based term frequency-inverse document frequency (TF-IDF) algorithms, among which TF-IDF is a commonly used weighting technique for information retrieval and data mining. This application does not limit the types of text similarity algorithms used.

[0075] It should also be understood that the terminal device uses a text similarity algorithm to calculate the user interest points corresponding to each of at least one preset vectorized data cluster.

[0076] For example, at least one preset vectorized data cluster corresponds to an interest point including: cluster number cluster_id, interest point interset_point{0, (game, Honor of Kings, PUBG); 1, (stock, fund, wealth management)}.

[0077] S306, based on the interest points corresponding to at least one preset vectorized data cluster, determine multiple preset interest points corresponding to multiple preset vectorized data as interest points of the preset vectorized data clusters corresponding to the multiple preset vectorized data.

[0078] S307, Obtain target usage data of the terminal device during the target time period. The target usage data includes usage data of at least one APP during the target time period, description data of at least one APP, and status data including screen-on information and screen-off information of the terminal device.

[0079] This step and Figure 2 Step S201 in the illustrated embodiment is similar and will not be repeated here.

[0080] S308, the target usage data is vectorized, and the similarity between the vectorized target usage data and multiple preset vectorized data is calculated to determine the target vectorized data. The target vectorized data is the vectorized data with the highest similarity to the vectorized target usage data among the multiple preset vectorized data. The multiple preset vectorized data are obtained by vectorizing historical usage data before the target time period.

[0081] This step and Figure 2 Step S202 in the illustrated embodiment is similar and will not be repeated here.

[0082] S309, determine the preset interest points corresponding to the target vectorized data as the target interest points for the target time period, and the multiple preset vectorized data correspond to the multiple preset interest points respectively.

[0083] This step and Figure 2 Step S203 in the illustrated embodiment is similar and will not be repeated here.

[0084] For example, taking at least one preset vectorized data cluster corresponding to points of interest including: cluster_id,interset_point{0,(game, Honor of Kings, PUBG); 1,(stock, fund, wealth management)} as an example, the target points of interest for the target time period may be session_id, cluster_id, and points of interest interset_point{0, 0,(game, Honor of Kings, PUBG)}.

[0085] This embodiment mainly illustrates how a terminal device determines multiple preset vectorized data and multiple preset points of interest corresponding to the multiple preset vectorized data based on historical usage data. This facilitates the terminal device to calculate the similarity between the vectorized target usage data and the multiple preset vectorized data, and then finally determine the target points of interest for the target time period. This method can greatly facilitate the accurate perception of a user's points of interest within a certain period of time without the user actively informing them, providing rich descriptions of user points of interest for downstream businesses, thereby improving the user experience.

[0086] In the above embodiments Figure 2 and Figure 3 Based on the embodiments shown, Figure 4 A flowchart of a method for extracting points of interest provided in another embodiment of this application is shown. Figure 4 In the illustrated embodiment, taking how the terminal device determines the target vectorized data as an example, the following is a detailed explanation. Figure 4 The flowchart illustrates the various steps in the method shown, including:

[0087] S401, acquire target usage data of the terminal device during a target time period. The target usage data includes usage data of at least one APP during the target time period, description data of at least one APP, and status data including screen-on information and screen-off information of the terminal device.

[0088] This step and Figure 2 Step S201 in the illustrated embodiment is similar and will not be repeated here.

[0089] S402, calculate the distance between the vectorized target data and the center coordinates of at least one preset vectorized data cluster.

[0090] It should be understood that the vectorized target data is a vector, and the terminal device calculates the distance between this vector and the center coordinates of at least one preset vectorized data cluster. For example, if there are 100 preset vectorized data clusters, then 100 distances are calculated.

[0091] S403, the preset vectorized data cluster corresponding to the minimum distance among the distances between the vectorized target data and the center coordinates of at least one preset vectorized data cluster is determined as the target vectorized data cluster.

[0092] It should be understood that the target vectorized data belongs to the preset vectorized data cluster corresponding to the minimum distance among the distances between the vectorized target data and the center coordinates of at least one preset vectorized data cluster.

[0093] S404, Based on the vectorized target usage data, determine the target vectorized data from the target vectorized data cluster.

[0094] It should be understood that the target vectorized data is the vectorized data in the target vectorized data cluster that has the highest similarity to the vectorized target data.

[0095] S405, determine the preset interest points corresponding to the target vectorized data as the target interest points for the target time period, and the multiple preset vectorized data correspond to the multiple preset interest points respectively.

[0096] This step and Figure 2 Step S203 in the illustrated embodiment is similar and will not be repeated here.

[0097] In the above technical solution, the terminal device determines which cluster the vectorized target usage data belongs to based on the distance between the center coordinates of the vectorized target usage data and at least one preset vectorized data cluster. This allows for the determination of the target vectorized data and the target interest points for the target time period. This method can greatly facilitate the accurate perception of a user's interest points within a certain time period without the user actively informing the user. It also facilitates the extraction of user interest points that change over time or periodically, while providing rich descriptions of user interest points for downstream businesses and improving user experience.

[0098] In one possible implementation, the above Figures 2 to 4 In the embodiments, the target usage data is obtained by filtering the initial usage data within the target time period using filtering rules, and the historical usage data is obtained by filtering the initial usage data within at least one historical time period using filtering rules. The initial usage data includes usage data of at least one APP, description data of at least one APP, and status data including screen-on information and screen-off information of the terminal device.

[0099] The filtering rules include one or more preset filtering rules, which include: filtering rules for filtering the usage data of at least one APP based on the opening and closing information of each APP; filtering rules for filtering status data based on the duration of screen on or off; and filtering rules for filtering the description data of at least one APP based on the number of downloads and the method of obtaining description data.

[0100] It should be understood that the filtering rules for filtering the usage data of at least one APP based on the opening and closing information of each APP include, but are not limited to, the following rules: if the time interval between opening and closing the first APP in at least one APP is less than a first preset time threshold, then the usage data corresponding to the first APP is deleted; if the duration after the first APP is opened is greater than a second preset time threshold, then the usage data corresponding to the first APP is deleted, wherein the second preset time threshold is greater than the first preset time threshold.

[0101] For example, if the first preset time threshold is 1 second and the second preset time threshold is 12 hours, then if the time interval between opening and closing the first APP in at least one APP is less than 1 second, the APP usage record is removed; if the first APP is continuously running for more than 12 hours, it is considered that the system recording function is abnormal, and the APP usage record is removed.

[0102] The 1 second and 12 hours in the above examples are just examples. These values ​​can be updated and adjusted according to user habits, etc., and this application does not limit them.

[0103] It should also be understood that the filtering rules for filtering the description data of at least one APP based on the number of downloads and the method of obtaining the description data include, but are not limited to, the following rules: if the total number of downloads of the second APP in at least one APP is less than a preset threshold, then the usage data and description data corresponding to the second APP are deleted; if the second APP does not have description data and the corresponding description data cannot be obtained, then the usage data and description data corresponding to the second APP are deleted.

[0104] For example, if the preset threshold is 10 times, then if the second app in at least one app is downloaded less than 10 times in the app store, all records of the second app are simultaneously removed from the app's description data and app usage data. Furthermore, after removal from the app's usage data, the specific character [SKP] is replaced. [SKP] data will be skipped and not included in the statistics in subsequent operations. For instance, in step S303 above, when "determining the session description text corresponding to each session based on at least one app in each session and the description data of at least one app, and obtaining multiple session description texts," data containing [SKP] records must be ignored.

[0105] It should be understood that if the second app has been downloaded less than 10 times in the app store, it means that the number of users of this app is too small and it is not of reference value, so the relevant record data has been deleted.

[0106] It should also be understood that if the description data of the second app is missing, then data can be populated from third-party websites using technologies such as web scraping.

[0107] If the description data for the second app is missing and cannot be filled in using other techniques, then the relevant records for the second app will be removed from both the app's description data and its usage data. After removal from the app's usage data, the specific character [SKP] will be replaced. The [SKP] data will be skipped and not included in subsequent statistics.

[0108] It should be understood that the filtering rules for filtering status data based on the duration of screen on or off include, but are not limited to, the following rules: if the time interval between the terminal device turning its screen on and off is less than a third preset time threshold, then the corresponding status data including the terminal device's screen on and off information is deleted; if the duration of the terminal device's screen on or off is greater than a fourth preset time threshold, then the corresponding status data including the terminal device's screen on and off information is deleted, wherein the fourth preset time threshold is greater than the third preset time threshold.

[0109] For example, if the third preset time threshold is 1 second and the second preset time threshold is 24 hours, then if the time interval between the terminal device turning off the screen and then turning on the screen is less than 1 second, then the corresponding status data including the screen-on information and screen-off information of the terminal device will be deleted; if the duration of the terminal device's screen-off or screen-on is greater than 24 hours, then the corresponding status data including the screen-on information and screen-off information of the terminal device will be deleted.

[0110] The 1 second and 24 hours in the above examples are just examples. These values ​​can be updated and adjusted according to user habits, etc., and this application does not limit them.

[0111] In summary, the interest point extraction method in this application utilizes text vectorization, text clustering, and text similarity techniques to extract user interest points over a certain period of time. This effectively solves the problem of complex functions in super apps, where user interest points cannot be accurately identified or displayed after identification. The solution in this application greatly facilitates the accurate perception of user interest points over a certain period of time without requiring the user to actively disclose them, making it easier to extract user interest points that change over time or periodically, while providing rich descriptions of user interest points for downstream businesses.

[0112] It should also be understood that the various embodiments described above can be coupled to each other, and this application does not limit this. Furthermore, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0113] The above text combines Figures 1 to 4 The method for extracting points of interest according to embodiments of this application is described in detail below. Figure 5 and Figure 6 This application describes in detail the point-of-interest extraction apparatus according to embodiments of the present application.

[0114] Figure 5 This is a structural schematic diagram of an interest point extraction device 500 provided in one embodiment of the present application. The device 500 includes: an acquisition module 501, a processing module 502, and a determination module 503.

[0115] The acquisition module 501 is used to acquire target usage data of a terminal device during a target time period. The target usage data includes usage data of at least one application (APP) during the target time period, description data of the at least one APP, and status data including screen-on and screen-off information of the terminal device. The processing module 502 is used to vectorize the target usage data and calculate the similarity between the vectorized target usage data and multiple preset vectorized data to determine target vectorized data. The target vectorized data is the vectorized data with the highest similarity to the vectorized target usage data among the multiple preset vectorized data. The multiple preset vectorized data are obtained by vectorizing historical usage data before the target time period. The determination module 503 is used to determine the preset interest point corresponding to the target vectorized data as the target interest point of the target time period. The multiple preset vectorized data correspond to multiple preset interest points respectively.

[0116] In some embodiments, the apparatus further includes a splitting module. The acquisition module 501 is further configured to acquire the historical usage data, which includes usage data of at least one APP within at least one historical time period, description data of the at least one APP, and status data including screen-on information and screen-off information of the terminal device; the splitting module is configured to split the historical usage data into multiple sub-data; and the processing module 502 is further configured to convert the multiple sub-data into vector representations to obtain the plurality of preset vectorized data.

[0117] In some embodiments, the splitting module is specifically used to: split the historical usage data into multiple sessions based on the status data, each session corresponding to a screen-on event or a screen-off event, each session including apps that are opened sequentially in the same screen-on event with an adjacent time interval of less than a preset time interval threshold; obtain sub-data corresponding to each session to obtain the multiple sub-data, each sub-data including usage data of at least one app included in the session and description data of at least one app.

[0118] In some embodiments, the processing module 502 is specifically configured to: determine the session description text corresponding to each session based on at least one APP in each session and the description data of the at least one APP, and obtain multiple session description texts; and use text vectorization technology to convert the multiple session description texts into vector representations to obtain the multiple preset vectorized data.

[0119] In some embodiments, the determining module 503 is further configured to determine multiple preset interest points corresponding to the multiple preset vectorized data.

[0120] In some embodiments, the determining module 503 is specifically configured to: merge the plurality of preset vectorized data according to a text clustering algorithm to obtain at least one preset vectorized data cluster; determine the interest points corresponding to the at least one preset vectorized data cluster based on a text similarity algorithm and the session description text corresponding to the at least one preset vectorized data cluster; and determine the plurality of preset interest points corresponding to the plurality of preset vectorized data as interest points of the preset vectorized data cluster corresponding to the plurality of preset vectorized data according to the interest points corresponding to the at least one preset vectorized data cluster.

[0121] In some embodiments, the target usage data is obtained by filtering the initial usage data within the target time period using filtering rules, and the historical usage data is obtained by filtering the initial usage data within the at least one historical time period using the filtering rules. The initial usage data includes usage data of at least one APP, description data of the at least one APP, and status data including screen-on information and screen-off information of the terminal device.

[0122] In some embodiments, the filtering rules include one or more preset filtering rules, which include: filtering rules for filtering the usage data of the at least one APP based on the opening and closing information of each APP; filtering rules for filtering the status data based on the duration of screen on or off; and filtering rules for filtering the description data of the at least one APP based on the number of downloads and the method of obtaining the description data.

[0123] In some embodiments, the device further includes a filtering module. The filtering module is specifically configured to: delete usage data corresponding to the first app if the time interval between opening and closing the first app in the at least one app is less than a first preset time threshold; and delete usage data corresponding to the first app if the duration after the first app is opened is greater than a second preset time threshold, wherein the second preset time threshold is greater than the first preset time threshold.

[0124] In some embodiments, the filtering module is specifically used to: if the total number of downloads of the second APP in the at least one APP is less than a preset number threshold, then delete the usage data and description data corresponding to the second APP; if the second APP does not have description data and the corresponding description data cannot be obtained, then delete the usage data and description data corresponding to the second APP.

[0125] In some embodiments, the filtering module is specifically used to: if the time interval between the terminal device turning off the screen and then turning on the screen is less than a third preset time threshold, then delete the corresponding state data including the screen-on information and screen-off information of the terminal device; if the duration of the terminal device's screen-off or screen-on is greater than a fourth preset time threshold, then delete the corresponding state data including the screen-on information and screen-off information of the terminal device, wherein the fourth preset time threshold is greater than the third preset time threshold.

[0126] In some embodiments, the processing module 502 is specifically configured to: calculate the distance between the vectorized target usage data and the center coordinates of the at least one preset vectorized data cluster; determine the preset vectorized data cluster corresponding to the minimum distance among the distances between the vectorized target usage data and the center coordinates of the at least one preset vectorized data cluster as the target vectorized data cluster; and determine the target vectorized data from the target vectorized data cluster based on the vectorized target usage data.

[0127] It should be understood that the device 500 here is embodied in the form of a functional module. The term "module" here can refer to application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors, etc.) and memories for executing one or more software or firmware programs, integrated logic circuits, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the device 500 can be specifically the terminal device in the above embodiments, or the functions of the terminal device in the above embodiments can be integrated into the device 500. The device 500 can be used to execute the various processes and / or steps corresponding to the terminal device in the above method embodiments; to avoid repetition, these will not be described further here.

[0128] The aforementioned device 500 has the function of implementing the corresponding steps performed by the terminal device in the aforementioned method; the aforementioned function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned function.

[0129] Figure 6 This is a structural schematic diagram of an apparatus provided for another embodiment of this application. Figure 6 The apparatus shown can be used to perform the method of any of the foregoing embodiments.

[0130] like Figure 6As shown, the device 600 in this embodiment includes a memory 601, a processor 602, a communication interface 603, and a bus 604. The memory 601, processor 602, and communication interface 603 are interconnected via the bus 604.

[0131] The memory 601 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 601 may store a program, and when the program stored in the memory 601 is executed by the processor 602, the processor 602 performs the various steps of the method shown in the above embodiments.

[0132] The processor 602 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, used to execute relevant programs to implement the various methods shown in the embodiments of this application.

[0133] The processor 602 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the method in this embodiment can be accomplished through integrated logic circuits in the processor 602 or through software instructions.

[0134] The processor 602 described above can also be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.

[0135] The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 601. The processor 602 reads the information in memory 601 and, in conjunction with its hardware, completes the functions required by the units included in the device of this application.

[0136] The communication interface 603 can use, but is not limited to, transceivers to enable communication between the device 600 and other devices or communication networks.

[0137] Bus 604 may include a pathway for transmitting information between various components of device 600 (e.g., memory 601, processor 602, communication interface 603).

[0138] It should be understood that the device 600 shown in the embodiments of this application may be an electronic device, or it may be a chip configured in an electronic device.

[0139] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0140] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0141] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0144] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0145] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0146] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for extracting points of interest, characterized in that, include: Acquire target usage data of a terminal device during a target time period. The target usage data includes usage data of at least one application (APP) during the target time period, description data of the at least one APP, and status data including screen-on information and screen-off information of the terminal device. The target usage data is vectorized, and the similarity between the vectorized target usage data and multiple preset vectorized data is calculated to determine the target vectorized data. The target vectorized data is the vectorized data with the highest similarity to the vectorized target usage data among the multiple preset vectorized data. The multiple preset vectorized data are obtained by vectorizing historical usage data before the target time period. The preset points of interest corresponding to the target vectorized data are determined as the target points of interest for the target time period, and the multiple preset vectorized data correspond to the multiple preset points of interest respectively.

2. The method according to claim 1, characterized in that, The method further includes: The historical usage data is obtained, which includes usage data of at least one APP within at least one historical time period, description data of the at least one APP, and status data including screen-on information and screen-off information of the terminal device. The historical usage data is split into multiple sub-data; The multiple sub-data are converted into vector representations to obtain the multiple preset vectorized data.

3. The method according to claim 2, characterized in that, The process of splitting the historical usage data into multiple sub-data sets includes: Based on the status data, the historical usage data is divided into multiple sessions, each session corresponding to a screen-on event or a screen-off event. Each session includes apps that are opened sequentially in the same screen-on event with an adjacent time interval of less than a preset time interval threshold. Obtain sub-data corresponding to each session to obtain the multiple sub-data sets. Each sub-data set includes usage data of at least one APP included in the session and description data of at least one APP.

4. The method according to claim 3, characterized in that, The step of converting the multiple sub-data into vector representations to obtain the multiple preset vectorized data includes: Based on at least one app in each session and the description data of the at least one app, determine the session description text corresponding to each session to obtain multiple session description texts; Using text vectorization technology, the multiple session description texts are converted into vector representations to obtain the multiple preset vectorized data.

5. The method according to claim 2, characterized in that, After converting the multiple sub-data into vector representations to obtain the multiple preset vectorized data, the method further includes: Determine multiple preset points of interest corresponding to the multiple preset vectorized data.

6. The method according to claim 5, characterized in that, The step of determining the multiple preset interest points corresponding to the multiple preset vectorized data includes: According to the text clustering algorithm, the multiple preset vectorized data are merged to obtain at least one preset vectorized data cluster; Based on the text similarity algorithm and the conversation description text corresponding to the at least one preset vectorized data cluster, the interest points corresponding to the at least one preset vectorized data cluster are determined. Based on the interest points corresponding to the at least one preset vectorized data cluster, the multiple preset interest points corresponding to the multiple preset vectorized data are determined as the interest points of the preset vectorized data cluster corresponding to the multiple preset vectorized data.

7. The method according to claim 2, characterized in that, The target usage data is obtained by filtering the initial usage data within the target time period using filtering rules. The historical usage data is obtained by filtering the initial usage data within the at least one historical time period using the filtering rules. The initial usage data includes usage data of at least one APP, description data of the at least one APP, and status data including screen-on information and screen-off information of the terminal device.

8. The method according to claim 7, characterized in that, The filtering rules include one or more preset filtering rules, which include: Filtering rules for filtering the usage data of at least one app based on the opening and closing information of each app; The filtering rules for filtering the status data based on the duration of screen on or off; The filtering rules for filtering the description data of at least one APP based on the number of downloads and the method of obtaining the description data.

9. The method according to claim 8, characterized in that, The filtering rules for filtering the usage data of at least one app based on the opening and closing information of each app include: If the time interval between opening and closing the first app in the at least one app is less than a first preset time threshold, then the usage data corresponding to the first app is deleted. If the first app remains open for a period of time longer than a second preset time threshold, then the usage data corresponding to the first app is deleted. The second preset time threshold is greater than the first preset time threshold.

10. The method according to claim 8, characterized in that, The filtering rules for filtering the description data of the at least one APP based on the number of downloads and the method of obtaining the description data include: If the total number of downloads of the second app in the at least one app is less than a preset threshold, then the usage data and description data corresponding to the second app are deleted. If the second app does not have description data and the corresponding description data cannot be obtained, then delete the usage data and description data corresponding to the second app.

11. The method according to claim 8, characterized in that, The filtering rules for filtering the status data based on the duration of screen on or off include: If the time interval between the screen being off and then on again is less than a third preset time threshold, then the corresponding status data including the screen-on information and screen-off information of the terminal device is deleted. If the duration of the terminal device's screen being off or on exceeds a fourth preset time threshold, then the corresponding status data including the terminal device's screen-on and screen-off information is deleted, where the fourth preset time threshold is greater than the third preset time threshold.

12. The method according to claim 6, characterized in that, The step of vectorizing the target usage data and calculating the similarity between the vectorized target usage data and multiple preset vectorized data to determine the target vectorized data includes: Calculate the distance between the vectorized target data and the center coordinates of the at least one preset vectorized data cluster; The preset vectorized data cluster corresponding to the minimum distance among the distances between the vectorized target data and the center coordinates of at least one preset vectorized data cluster is determined as the target vectorized data cluster; Based on the vectorized target usage data, the target vectorized data is determined from the target vectorized data cluster.

13. A device for extracting points of interest, characterized in that, include: The acquisition module is used to acquire target usage data of the terminal device during a target time period. The target usage data includes usage data of at least one application (APP) during the target time period, description data of the at least one APP, and status data including screen-on information and screen-off information of the terminal device. The processing module is used to vectorize the target usage data and calculate the similarity between the vectorized target usage data and multiple preset vectorized data to determine the target vectorized data. The target vectorized data is the vectorized data with the highest similarity to the vectorized target usage data among the multiple preset vectorized data. The multiple preset vectorized data are obtained by vectorizing historical usage data before the target time period. The determination module is used to determine the preset interest points corresponding to the target vectorized data as the target interest points of the target time period, wherein the plurality of preset vectorized data correspond to the plurality of preset interest points respectively.

14. A terminal device, characterized in that, It includes a processor and a memory, the memory being used to store code instructions; the processor being used to execute the code instructions to perform the method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, Used to store a computer program, the computer program including instructions for implementing the method as described in any one of claims 1 to 12.