Information data processing methods, apparatus, equipment and media

By calculating the number of hot words and the frequency of occurrence of information data to determine its value parameters, the problem of being unable to accurately search for high-value information data in the information explosion era is solved, and the accurate display of high-value information data is achieved, thus improving the user experience.

CN118862866BActive Publication Date: 2025-10-28RICHFIT INFORMATION TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310465937.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-10-28
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

In the information explosion era, searching for information based on keywords cannot accurately retrieve high-value information, resulting in poor search results.

Method used

By collecting information data, we determine the number of hot words and the frequency of their occurrence. Based on these indicators, we calculate the value parameters of the information data and determine the display order or filter information data according to the value parameters to ensure that high-value information data is displayed to users.

Benefits of technology

It improved the accuracy of information and data display, reduced the amount of low-value information and data that users read, and improved user research efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118862866B_ABST
    Figure CN118862866B_ABST
Patent Text Reader

Abstract

This application provides an information data processing method, apparatus, device, and medium, belonging to the field of computer technology. The method includes: collecting information data belonging to a first domain; determining the number of hot words and the frequency of their occurrence based on reference data corresponding to the first domain, wherein the reference data includes multiple hot words corresponding to the first domain; determining the value parameters of the information data based on the number of hot words and their frequency of occurrence; and displaying the information data based on the value parameters. This solution ensures the accuracy of the determined value parameters of the information data and improves the accuracy of the information data displayed to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an information data processing method, apparatus, device, and medium. Background Technology

[0002] When users want to research a specific topic, they can enter keywords on search engines to find relevant information and data, thus acquiring knowledge in that field. However, in this era of information overload, the internet contains a vast amount of low-value information. Searching solely based on keywords will only yield results containing those keywords, failing to accurately retrieve high-value information, forcing users to read excessive amounts of data. Therefore, the accuracy of the information search results obtained using the aforementioned method is poor. Summary of the Invention

[0003] This application provides an information data processing method, apparatus, device, and medium, which ensures the accuracy of the determined value parameters of the information data and improves the accuracy of the information data displayed to the user. The technical solution is as follows:

[0004] On the one hand, an information data processing method is provided, the method comprising:

[0005] Collect information data belonging to the first domain;

[0006] Based on the reference data corresponding to the first domain, the number of hot words and the number of times hot words appear in the information data are determined. The reference data includes multiple hot words corresponding to the first domain. The number of hot words in the information data is the number of hot words that exist in the information data among the multiple hot words corresponding to the first domain. The number of times hot words appear is the number of times hot words appear in the information data.

[0007] The value parameters of the information data are determined based on the number of hot words and the frequency of their occurrence.

[0008] The information data is displayed based on its value parameters.

[0009] In one possible implementation, determining the value parameters of the information data based on the number of hot words and the frequency of their occurrence includes:

[0010] Keyword extraction is performed on the information data to obtain the keywords of the information data;

[0011] The value parameters of the information data are determined based on the number of hot words, the number of times the hot words appear, and the number of times the keywords appear in the information data.

[0012] In one possible implementation, determining the value parameters of the information data based on the number of hot words and the frequency of their occurrence includes:

[0013] The value parameters of the information data in multiple dimensions are determined, including domain relevance, which is determined based on the number of hot words and the frequency of occurrence of hot words in the information data.

[0014] In one possible implementation, the value parameters of the multiple dimensions also include at least one of information credibility, content maturity, information dissemination, information attention, or information search popularity.

[0015] In one possible implementation, the credibility of the information is determined based on at least one of the data source of the information data and the number of characters in the body of the information data; or...

[0016] The content maturity level is determined based on at least one of the following: the frequency of occurrence of keywords in the information data and the relevance between the semantics of the title and the semantics of the body text of the information data; or...

[0017] The dissemination of the information is determined based on at least one of the following: the number of media outlets publishing the information data, the number of times the information data is shared, and the number of times the information data is forwarded; or...

[0018] The level of attention given to the information is determined based on at least one of the number of times the information data is saved and the number of times it is liked; or...

[0019] The search popularity of the information is determined based on the search information corresponding to the keywords in the information data.

[0020] In one possible implementation, the reference data further includes multiple data sources and the credibility of the multiple data sources; determining the credibility of the information based on at least one of the data source of the information data and the word count of the main body of the information data includes:

[0021] Obtain the first credibility level corresponding to the data source of the information data from the reference data;

[0022] Based on the number of characters in the main body of the information data, a second credibility of the information data is determined, and the second credibility is positively correlated with the number of characters in the main body;

[0023] The credibility of the information is determined based on at least one of the first credibility and the second credibility.

[0024] In one possible implementation, the method further includes:

[0025] Collect multiple pieces of information data belonging to the first domain;

[0026] Keyword extraction is performed on the multiple pieces of information data to obtain the keywords of the multiple pieces of information data;

[0027] The keywords of the multiple pieces of information data are used as hot words corresponding to the first domain; or, the keywords of the multiple pieces of information data that appear at least once a first threshold number of times are used as hot words corresponding to the first domain.

[0028] In one possible implementation, the method further includes:

[0029] Display a data configuration interface, which is used to obtain reference data corresponding to the first domain;

[0030] The data entered in the data configuration interface is used as the reference data corresponding to the first domain.

[0031] In one possible implementation, displaying the information data based on the value parameters of the information data includes:

[0032] The multiple pieces of information data are displayed in descending order of their value parameters; or,

[0033] The information data is displayed if the value parameter of the information data is not less than the second threshold.

[0034] On the other hand, an information data processing apparatus is provided, the apparatus comprising:

[0035] The data acquisition module is used to collect information data belonging to the first domain.

[0036] The first determining module is used to determine the number of hot words and the number of times hot words appear in the information data based on the reference data corresponding to the first domain. The reference data includes multiple hot words corresponding to the first domain. The number of hot words in the information data is the number of hot words that exist in the information data among the multiple hot words corresponding to the first domain. The number of times hot words appear is the number of times hot words appear in the information data.

[0037] The second determining module is used to determine the value parameters of the information data based on the number of hot words and the frequency of hot word occurrences in the information data;

[0038] The display module is used to display the information data based on the value parameters of the information data.

[0039] In one possible implementation, the second determining module includes:

[0040] The keyword extraction unit is used to extract keywords from the information data to obtain the keywords of the information data;

[0041] The value parameter determination unit is used to determine the value parameters of the information data based on the number of hot words, the number of times the hot words appear, and the number of times the keywords appear in the information data.

[0042] In one possible implementation, the second determining module is used to determine the value parameters of the information data in multiple dimensions, including domain relevance, which is determined based on the number of hot words and the frequency of occurrence of hot words in the information data.

[0043] In one possible implementation, the value parameters of the multiple dimensions also include at least one of information credibility, content maturity, information dissemination, information attention, or information search popularity.

[0044] In one possible implementation, the credibility of the information is determined based on at least one of the data source of the information data and the number of characters in the body of the information data; or...

[0045] The content maturity level is determined based on at least one of the following: the frequency of occurrence of keywords in the information data and the relevance between the semantics of the title and the semantics of the body text of the information data; or...

[0046] The dissemination of the information is determined based on at least one of the following: the number of media outlets publishing the information data, the number of times the information data is shared, and the number of times the information data is forwarded; or...

[0047] The level of attention given to the information is determined based on at least one of the number of times the information data is saved and the number of times it is liked; or...

[0048] The search popularity of the information is determined based on the search information corresponding to the keywords in the information data.

[0049] In one possible implementation, the reference data further includes multiple data sources and the credibility of the multiple data sources; the second determining module includes:

[0050] A value parameter determination unit is used to obtain a first credibility level corresponding to the data source of the information data from the reference data; determine a second credibility level of the information data based on the number of characters in the main body of the information data, wherein the second credibility level is positively correlated with the number of characters in the main body; and determine the credibility level of the information based on at least one of the first credibility level and the second credibility level.

[0051] In one possible implementation, the device further includes:

[0052] The acquisition module is also used to acquire multiple pieces of information data belonging to the first domain;

[0053] The keyword extraction module is used to extract keywords from the multiple pieces of information data to obtain the keywords of the multiple pieces of information data;

[0054] The hot word determination module is used to select the keywords of the multiple pieces of information data as hot words corresponding to the first domain; or, select the keywords of the multiple pieces of information data whose occurrence frequency is not less than a first threshold as hot words corresponding to the first domain.

[0055] In one possible implementation, the device further includes:

[0056] The display module is used to display a data configuration interface, which is used to obtain reference data corresponding to the first domain.

[0057] The reference data acquisition module is used to use the data input in the data configuration interface as the reference data corresponding to the first domain.

[0058] In one possible implementation, the display module is used to display the multiple pieces of information data in descending order of their value parameters; or,

[0059] The display module is used to display the information data when the value parameter of the information data is not less than a second threshold.

[0060] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the information data processing method as described in any of the above implementations.

[0061] On the other hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the at least one piece of program code being loaded and executed by a processor to implement the information data processing method as described in any of the above implementations.

[0062] On the other hand, a computer program product is provided, the computer program product including at least one piece of program code, the at least one piece of program code being loaded and executed by a processor to implement the information data processing method as described in any of the above implementations.

[0063] The beneficial effects of the technical solutions provided by the embodiments of the present application include at least:

[0064] This application provides an information data processing method. After collecting information data, the method determines the value parameters of the information data. These value parameters are determined based on the number of hot words belonging to the relevant field and the frequency of their occurrence, ensuring the accuracy of the determined value parameters. Based on these value parameters, the information data is displayed, enabling the presentation of high-value information data to users and improving the accuracy of the information data presented. Attached Figure Description

[0065] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0066] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0067] Figure 2 This is a flowchart of an information data processing method provided in an embodiment of this application;

[0068] Figure 3 This is a flowchart of an information data processing method provided in an embodiment of this application;

[0069] Figure 4 This is a schematic diagram of the structure of an information data processing device provided in an embodiment of this application;

[0070] Figure 5 This is a schematic diagram of the structure of an information data processing device provided in an embodiment of this application;

[0071] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0072] Figure 7 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0074] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0075] The information data processing method provided in this application is executed by a computer device. In some embodiments, the computer device may be an electronic device, such as a mobile phone, tablet computer, or desktop computer; this application does not limit the type of electronic device. In other embodiments, the computer device may be a server, which may be a single server, a server cluster consisting of several servers, or a cloud computing service center. Of course, the server may also include other functional servers to provide more comprehensive and diversified services. In other embodiments, the computer device includes both electronic devices and a server. It should be noted that this application is merely an illustrative description of the executing entity of the information data processing method and does not limit the executing entity.

[0076] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application, such as... Figure 1 As shown, the implementation environment includes electronic device 101 and server 102, which are connected via wired or wireless network.

[0077] In some embodiments, server 102 collects information data belonging to a first domain from the network, determines the value parameters of the information data, and sends the information data to electronic device 101 when the value parameters of the information data meet the conditions, so that electronic device 101 can display the information data; or, after determining the value parameters of the information data, server 102 sends the information data and the corresponding value parameters to electronic device 101, so that electronic device 101 can display the information data based on the value parameters of the information data.

[0078] Figure 2 This is a flowchart illustrating an information data processing method provided in an embodiment of this application. The embodiment uses a computer device as an example for illustrative purposes. See also... Figure 2 The method includes:

[0079] 201. Computer equipment collects information data belonging to the first domain.

[0080] The first field can be any field, such as oil extraction, software development, plastic manufacturing, etc. The embodiments of this application do not limit the first field.

[0081] Information data refers to information that users can obtain promptly and that can bring value to users within a relatively short period of time. In some embodiments, the information data can be any form of data transmitted over the Internet, such as papers published in journals, articles published on public accounts, or user blog posts. In some embodiments, the information data includes at least one of text data, image data, video data, and audio data. This application embodiment does not limit the content style of the information data.

[0082] In some embodiments, a computer device can search for information data from the network based on domain terms corresponding to a first domain. The computer device collecting information data belonging to the first domain includes: the computer device collecting information data matching the domain terms from the network. Domain terms corresponding to the first domain are frequently used terms in research within that domain; they can be research objects within that domain, or processing techniques within that domain, etc. This application embodiment does not limit the scope of the domain terms. Optionally, the domain terms corresponding to the first domain are set by the user. Optionally, the domain terms corresponding to the first domain are determined based on information data belonging to the first domain; this application embodiment does not limit the method of determining the domain terms corresponding to the first domain.

[0083] In some embodiments, the computer device can collect information data from the entire network. In other embodiments, the computer device can collect information data from multiple preset data sources. Collecting information data belonging to a first domain includes: the computer device collecting information data corresponding to the first domain from multiple preset data sources. The multiple preset data sources can be data sources corresponding to the first domain or data sources specified by the user; this application embodiment does not limit the number of preset data sources. In some embodiments, data sources may include the Internet, public accounts, e-journals, search engines, and professional data sources, etc.; this application embodiment does not limit the data source for information data.

[0084] 202. The computer device determines the number of hot words and the number of times hot words appear in the information data based on the reference data corresponding to the first domain. The reference data includes multiple hot words corresponding to the first domain. The number of hot words in the information data is the number of hot words in the information data that exist among the multiple hot words corresponding to the first domain. The number of times hot words appear is the number of times each hot word appears in the information data.

[0085] In this embodiment, the hot words corresponding to the first domain refer to words that are highly relevant to the first domain. In some embodiments, the hot words of the first domain can be obtained by extracting keywords from information data belonging to the first domain. Of course, the hot words of the first domain can also be obtained in other ways, such as by user specification, etc. This embodiment does not limit the method of obtaining the hot words of the first domain.

[0086] The number of hot words in the news data is the number of hot words that exist in the news data among the multiple hot words corresponding to the first domain. For example, the multiple hot words corresponding to the first domain are "reservoir", "well logging", and "separation". If the news data contains "reservoir" and "well logging", then the number of hot words in the news data is 2. The number of times the hot words appear is the number of times "reservoir" appears in the news data and the number of times "well logging" appears in the news data, or the number of times the hot words appear is the sum of the number of times "reservoir" appears in the news data and the number of times "well logging" appears in the news data.

[0087] 203. Computer equipment determines the value parameters of information data based on the number of hot words and the frequency of their occurrence.

[0088] If multiple hot keywords corresponding to the first domain appear in a piece of information data, and these hot keywords also appear frequently, it indicates a high relevance of the information data to the first domain, and therefore, its value is high. This also demonstrates that the number and frequency of hot keywords in information data can accurately determine its value parameters.

[0089] 204. Computer equipment displays information data based on the value parameters of the information data.

[0090] Computer devices display information data based on its value parameters, allowing users to see high-value information instead of reading large amounts of low-value data. This improves both the accuracy of the information and the efficiency of user research. In some embodiments, the computer device can display information data in descending order of value parameters. That is, information with higher value parameters is placed first, followed by information with lower value parameters, ensuring users can access high-value information first and improving the user experience.

[0091] In some embodiments, the computer device displays information data based on a value parameter, including: displaying the information data when the value parameter is not less than a first threshold. That is, only high-value information data is displayed, while low-value information data is not, so that users can read high-value information data, thus improving the user experience.

[0092] It should be noted that the embodiments of this application are merely examples of the two display methods described above to illustrate the value parameters based on information data and the display of information data, and are not intended to limit the scope of the application. The method of displaying information data based on the value parameters of information data can be set according to the actual application needs.

[0093] The information data processing method provided in this application determines the value parameters of the information data after collecting it. These value parameters are determined based on the number of hot words belonging to the field and the frequency of their occurrence, ensuring the accuracy of the determined value parameters. By displaying the information data based on these value parameters, high-value information data can be presented to users, improving the accuracy of the information data presented to them.

[0094] Figure 3 This is a flowchart illustrating an information data processing method provided in an embodiment of this application. This embodiment uses a computer device as an example for illustrative purposes. See also... Figure 3 The method includes:

[0095] 301. Computer equipment collects information data belonging to the first domain.

[0096] In some embodiments, a user triggers a computer device to collect information data belonging to a first field by entering the name of that field in a search page. For example, a user entering "oil extraction" in the computer device's search page will cause the computer device to collect information data belonging to the oil extraction field.

[0097] In other embodiments, users do not need to enter keywords on the search page; the computer device can automatically collect information data belonging to the first domain. For example, to facilitate research on the first domain, an application or webpage corresponding to the first domain can be developed. The computer device is a device that captures information data for the application or webpage. The computer device continuously collects information data belonging to the first domain from the network and provides the information data to the application or webpage.

[0098] 302. The computer device acquires reference data corresponding to the first domain, the reference data including at least one of multiple data sources, the credibility of the multiple data sources, domain terms corresponding to the first domain, and multiple hot terms corresponding to the first domain.

[0099] In some embodiments, the reference data includes multiple data sources but does not include the credibility levels corresponding to the multiple data sources. When the data source of the information data collected by the computer device is the same as any of the data sources included in the reference data, the data source of the information data is determined to be credible; when the data source of the information data collected by the computer device is different from each of the data sources included in the reference data, the data source of the information data is determined to be untrustworthy. In other embodiments, the reference data includes multiple data sources and the credibility levels corresponding to the multiple data sources. When the data source of the information data collected by the computer device is the same as any of the data sources included in the reference data, the credibility level corresponding to that data source in the reference data can be used as the credibility level of the data source of the information data, or the credibility level corresponding to that data source can be used as the credibility level of the information data.

[0100] In some embodiments, the reference data includes domain terms corresponding to a first domain. The computer device performs step 301 above based on the domain terms in the reference data to collect information data matching the domain terms.

[0101] In some embodiments, the reference data includes hot words corresponding to a first domain. Hot words corresponding to the first domain are words that appear multiple times in the information data belonging to the first domain.

[0102] In one possible implementation, the reference data is user-configured. The computer device acquires the reference data corresponding to the first domain by: displaying a data configuration interface for acquiring the reference data corresponding to the first domain; and using the data input in the data configuration interface as the reference data corresponding to the first domain.

[0103] Optionally, the data configuration interface includes a first input option for inputting a data source, and the computer device determines the data obtained from the first input option as the data source. Optionally, the data configuration interface includes a first input option for inputting a data source and a second input option for inputting the credibility of the corresponding data source, and the computer device determines the data obtained from the first input option as the data source and the data obtained from the second input option as the credibility of the corresponding data source. The first input option and the second input option correspond one-to-one. In some embodiments, the data configuration interface can display multiple first input options and multiple second input options, allowing the user to input multiple data sources and their corresponding credibility through these multiple first input options and multiple second input options. In other embodiments, the data configuration interface can display one first input option and one second input option, allowing the user to input multiple data sources and their corresponding credibility through multiple inputs.

[0104] Optionally, the data configuration interface includes a third input option for inputting domain terms, and the computer device determines the data obtained from the third input option as the domain terms of the first domain.

[0105] Optionally, the data configuration interface includes a fourth input option for inputting hot words corresponding to the first domain, and the computer device determines the data obtained from the fourth input option as the hot words corresponding to the first domain.

[0106] In another possible implementation, the multiple hot words corresponding to the first domain are obtained by a computer device through summarizing information data of the first domain. In some embodiments, the process of obtaining multiple hot words corresponding to the first domain may include: collecting multiple pieces of information data belonging to the first domain; extracting keywords from the multiple pieces of information data to obtain keywords for the multiple pieces of information data; using the keywords of the multiple pieces of information data as hot words corresponding to the first domain, or using keywords from the multiple pieces of information data whose frequency of occurrence is not less than a first threshold as hot words corresponding to the first domain. The first threshold can be any value, such as 3, 5, 7, etc. Optionally, the first threshold is determined based on empirical values; alternatively, the first threshold is set by a technician, and this application embodiment does not limit the first threshold.

[0107] The keywords of the information data can be determined semantically or through a keyword extraction model. This application embodiment does not limit this, but is only for illustrative purposes. Optionally, the computer device extracts keywords from the information data to obtain the keywords of the information data, including: performing word segmentation on the information data to obtain multiple word segmentation results; obtaining the semantic features of each word segmentation result; and determining the word segmentation result whose semantic features are most similar to the semantic features of the information data as the keywords of the information data.

[0108] Optionally, the computer device performs keyword extraction on the information data to obtain the keywords of the information data, including: extracting keywords from the information data through a keyword extraction model to obtain the keywords of the information data, wherein the keyword extraction model is used to extract keywords from the information data.

[0109] The training process of the keyword extraction model will be illustrated below:

[0110] In some embodiments, a computer device collects multiple pieces of information data belonging to a first domain, establishes an information database containing these multiple pieces of information data, and performs data cleaning and preprocessing on the information database to obtain a corpus. In some embodiments, data cleaning may include at least one of deduplication, removal of sensitive words, and handling of missing values. Sensitive words can be any word set by the user; this application embodiment does not limit the types of sensitive words. Removing sensitive words may include: replacing sensitive words in the information data with preset words; or deleting the information data containing sensitive words.

[0111] For example, to prevent information data that we do not have permission to process, we can set some sensitive words and delete such information data based on these sensitive words.

[0112] In some embodiments, missing values ​​can be handled by padding the missing information data with 0. This application does not limit the method of handling missing values.

[0113] In some embodiments, preprocessing the information database may include format conversion of the information data to unify the format and standardize the information data. For example, some information data contains dates in the format xxx year xxx month xxx day, some in the format xxx / xxx / xxx, and some in the format xxx-xxx-xxx. A computer device can convert the dates in all information data in the database to the xxx-xxx-xxx format, thus unifying the date format of different information data in the database.

[0114] In some embodiments, preprocessing the information data may include normalizing the information data. For example, some information data is relatively short, while some is relatively long; therefore, the number of characters in the main body of the information data may vary significantly. This large difference in the number of characters may affect the accuracy of the value parameters when determining the value parameters of the information data. Normalizing the information data may involve normalizing the numerical format information, such as the number of characters in the main body of the information data.

[0115] It should be noted that information in the database is stored as a single piece of information data, while in the corpus, a piece of data can be a complete piece of information data or a part of information data. For example, a piece of data in the corpus is a paragraph of information data.

[0116] In some embodiments, the computer device selects at least a portion of the corpus from a corpus (which may be randomly selected or selected according to certain conditions, and this application embodiment does not limit this) to form a dataset. The user annotates the dataset with keywords, that is, the user annotates the keywords of any corpus in the dataset. The computer device divides the annotated dataset into multiple disjoint subsets, selects one subset as a test set, and uses the remaining subsets as a training set. The keyword extraction model is trained based on the training set. After training, the keyword extraction model is tested using the test set to obtain an evaluation metric. This process is repeated until each subset is selected as a training set. The computer device can determine the optimal keyword extraction model based on the evaluation metric obtained in each training round. This optimal keyword extraction model is the keyword extraction model corresponding to the highest evaluation metric obtained in each training round. Subsequently, the computer device can use the optimal keyword extraction model to extract keywords from the information database, obtain keywords from multiple pieces of information data in the information database, and use the keywords from these multiple pieces of information data as hot words corresponding to the first domain, or use the keywords from these multiple pieces of information data whose frequency of occurrence is not less than a first threshold as hot words corresponding to the first domain.

[0117] 303. Based on the reference data corresponding to the first domain, the computer equipment determines the value parameters of the information data in multiple dimensions.

[0118] In this embodiment, the value parameters of information data across multiple dimensions include at least one of the following: domain relevance, content maturity, information dissemination, information attention, or information search popularity. Specifically, domain relevance is the value parameter of information data in the domain relevance dimension, content maturity is the value parameter of information data in the content maturity dimension, information dissemination is the value parameter of information data in the dissemination dimension, information attention is the value parameter of information data in the attention dimension, and information search popularity is the value parameter of information data in the search dimension.

[0119] In one possible implementation, the value parameters of information data across multiple dimensions include domain relevance, which is the value parameter of information data in the domain relevance dimension.

[0120] In some embodiments, the reference data includes multiple hot words corresponding to a first domain. The computer device determines the domain relevance of the information data based on the number of hot words and the frequency of their occurrence. For example, the computer device determines the value parameters of the information data in multiple dimensions based on the reference data corresponding to the first domain, including: determining the number of hot words and the frequency of their occurrence based on the multiple hot words corresponding to the first domain, and determining the domain relevance of the information data based on the number of hot words and the frequency of their occurrence. The domain relevance is positively correlated with the number of hot words and the frequency of their occurrence.

[0121] In other embodiments, the reference data includes multiple domain terms and multiple hot words corresponding to the first domain. The computer device determines the domain relevance of the information data based on the number of hot words, the frequency of hot word occurrences, the number of domain terms, and the frequency of domain term occurrences. For example, the computer device determines the value parameters of the information data in multiple dimensions based on the reference data corresponding to the first domain, including: determining the number of hot words and the frequency of hot word occurrences based on the multiple hot words corresponding to the first domain; determining the number of domain terms and the frequency of domain term occurrences based on the multiple domain terms corresponding to the first domain; and determining the domain relevance of the information data based on the number of hot words, the frequency of hot word occurrences, the number of domain terms, and the frequency of domain term occurrences. Here, the number of domain terms in the information data is the number of domain terms present in the information data among the multiple domain terms corresponding to the first domain, and the frequency of domain term occurrences is the number of times the domain terms appear in the information data.

[0122] Optionally, the computer device collects multiple pieces of information data belonging to the first domain to form an information database. For each piece of information data in the information database, a value parameter is determined. Therefore, the value parameter of the information data can be used to represent the value of the information data relative to the value of other information data in the information database. Based on the number of hot words, the frequency of hot word occurrences, the number of domain terms, and the frequency of domain term occurrences, the computer device determines the domain relevance of the information data, including: normalizing the number of hot words to obtain a hot word count index; normalizing the frequency of hot word occurrences to obtain a hot word frequency index; normalizing the number of domain terms to obtain a domain term count index; normalizing the frequency of domain term occurrences to obtain a domain term frequency index; and weighting the hot word count index, hot word frequency index, domain term count index, and domain term frequency index to obtain the domain relevance of the resource data.

[0123] The weighting coefficients for the hot word count, hot word frequency, domain word count, and domain word frequency indicators can be configured by the user. This application embodiment does not limit the weighting coefficients for each indicator.

[0124] Similarly, computer equipment determines the domain relevance of information data based on the number of hot words and the frequency of hot word occurrences. This includes: normalizing the number of hot words to obtain a hot word count index; normalizing the frequency of hot words to obtain a hot word frequency index; and weighting the hot word count index and the hot word frequency index to obtain the domain relevance of the resource data.

[0125] In one possible implementation, the value parameters of the news data across multiple dimensions include content maturity. This content maturity is determined based on at least one of the following: the frequency of keyword occurrences in the news data and the relevance between the semantics of the news data's title and its body text. Specifically, keywords in the news data indicate its primary descriptive object. The multiple occurrences of keywords suggest that the news data has a clear theme and is therefore considered relatively mature.

[0126] The process by which computer equipment determines the content maturity of information data may include: extracting keywords from the information data to obtain its keywords and determining the frequency of each keyword's occurrence; extracting semantics from the title of the information data to obtain its title semantics; extracting semantics from the body text of the information data to obtain its body semantics; determining the relevance between the title semantics and the body semantics; and determining the content maturity of the information data based on the frequency of keyword occurrences and the relevance. Content maturity is positively correlated with both the frequency of keyword occurrences and the relevance.

[0127] In some embodiments, the computer device performs semantic extraction on the title of information data, including: extracting keywords from the title of the information data to obtain a first keyword, and obtaining the word vector of the first keyword, wherein the first keyword is a keyword corresponding to the title, and the word vector of the first keyword is used to represent the semantics of the title of the information data. The computer device performs semantic extraction on the body text of the information data, including: extracting keywords from the body text of the information data to obtain a second keyword, and obtaining the word vector of the second keyword, wherein the second keyword is a keyword corresponding to the body text, and the word vector of the second keyword is used to represent the semantics of the body text of the information data. The computer device determines the relevance between the title semantics and the body text semantics, including: determining the relevance between the word vectors of the first keyword and the word vectors of the second keyword. Optionally, the relevance is the cosine similarity between the word vectors of the first keyword and the word vectors of the second keyword. Optionally, the relevance is the Manhattan distance between the word vectors of the first keyword and the word vectors of the second keyword. Optionally, the relevance is the Pearson correlation coefficient between the word vectors of the first keyword and the word vectors of the second keyword. This application embodiment does not limit the relevance.

[0128] In some embodiments, the computer device determines the content maturity of information data based on the frequency of occurrence of the keyword and the relevance, including: the computer device normalizes the frequency of occurrence of the keyword to obtain a keyword frequency index of the information data; and weights the keyword frequency index with the relevance to obtain the content maturity of the information data.

[0129] The weighting coefficient between the keyword frequency index and the relevance can be any value, and this application does not limit this.

[0130] In one possible implementation, the value parameters of these multiple dimensions include information credibility. In some embodiments, the information credibility can be determined based on at least one of the data source of the information data and the word count of the information data. That is, the process by which the computer device determines the information credibility includes: the computer device determining a first credibility based on reference data and the data source of the information data; determining a second credibility of the information data based on the word count of the information data, wherein the second credibility is positively correlated with the word count; and determining the information credibility based on at least one of the first credibility and the second credibility.

[0131] Optionally, the reference data includes multiple reliable data sources. When the data source of the information data is the same as any data source in the reference data, it indicates that the data source of the information data is reliable. When the data source of the information data is not the same as any data source in the reference data, it indicates that the data source of the information data is unreliable. For example, a computer device determines a first confidence level based on the reference data and the data source of the information data, including: when the data source of the information data is the same as any data source in the reference data, determining the first confidence level as a first value; when the data source of the information data is not the same as any data source in the reference data, determining the first confidence level as a second value, wherein the second value is less than the first value. Optionally, the first value is 1 and the second value is 0. This application embodiment does not limit the first and second values.

[0132] Optionally, the reference data includes multiple data sources and the corresponding credibility levels of each data source. The computer device determines a first credibility level based on the reference data and the data source of the information data, including: obtaining the first credibility level corresponding to the data source of the information data from the reference data. This first credibility level can be considered an information source indicator for the information data.

[0133] Optionally, the computer device determines a second credibility of the information data based on the number of characters in the main body of the information data, including: normalizing the number of characters in the main body of the information data to obtain the second credibility of the information data, which can be regarded as the number of characters in the main body of the information data.

[0134] Optionally, the computer device determines the credibility of the information based on at least one of a first credibility level and a second credibility level, including: the computer device weighting the first credibility level and the second credibility level to obtain the information credibility level. The weighting coefficients for the first credibility level and the second credibility level can be any coefficients, and this embodiment of the application does not limit this.

[0135] In one possible implementation, the value parameters of these multiple dimensions include information dissemination, which is determined based on at least one of the number of media outlets publishing the information data, the number of times the information data is shared, and the number of times the information data is forwarded.

[0136] In some embodiments, the number of times information data is shared and the number of times information data is forwarded refer to the number of times the information data is shared and forwarded within the system after it is collected by the system provided in this application embodiment. Therefore, when the computer device first obtains the information data, it can determine the information dissemination degree of the information data based on the number of media outlets publishing the information data. Subsequently, the information dissemination degree of the information data can be updated at regular intervals based on at least one of the number of media outlets publishing the information data, the number of times the information data is shared, and the number of times the information data is forwarded.

[0137] In some embodiments, the process of a computer device determining the dissemination degree of information may include: normalizing the number of media outlets publishing the information data to obtain an index of the number of media outlets publishing the information data; normalizing the sum of the number of times the information data is shared and the number of times it is forwarded to obtain an index of the sharing of the information data; and weighting the index of the number of media outlets publishing the information data and the index of sharing to obtain the dissemination degree of the information data.

[0138] The weighting coefficients for the number of media outlets publishing the content and the sharing index can be any value, and this application does not limit this.

[0139] In one possible implementation, the multiple-dimensional value parameters include information attention. This information attention is determined based on at least one of the number of times the information data is saved and liked. In some embodiments, the number of times the information data is saved and liked refers to the number of times the information data is saved and liked within the system provided in this application embodiment after the system collects the information data. Therefore, when the system first obtains the information data, the information attention of the information data is zero. Subsequently, the system can update the information attention at regular intervals.

[0140] In some embodiments, the process of determining the information attention level of information data may include: normalizing the number of times the information data is collected to obtain the collection index of the information data; normalizing the number of times the information data is liked to obtain the liking index of the information data; and weighting the collection index and the liking index to obtain the information attention level of the information data.

[0141] The collection index and the liking index can be any value, and this application embodiment does not limit them.

[0142] In one possible implementation, the value parameters across multiple dimensions include information search popularity, which is determined based on search data corresponding to keywords in the information data. In some embodiments, search data may include search counts, changes in search counts, etc. In some embodiments, search data may include keyword search data in search engines such as Baidu, search data in mobile news applications such as Toutiao, etc., but this application does not limit the scope of search data.

[0143] It should be noted that the embodiments in this application are merely illustrative examples of the process for determining the value parameters of information data in multiple dimensions. In another embodiment, the value parameters of the information data can be determined directly based on the number of hot words and the frequency of their occurrence; alternatively, they can be determined based on the number of hot words, the frequency of their occurrence, and the frequency of keyword occurrences within the information data. For example, keywords can be extracted from the information data to obtain its keywords; the value parameters of the information data can then be determined based on the number of hot words, the frequency of their occurrence, and the frequency of keyword occurrences within the information data.

[0144] 304. Computer equipment determines the comprehensive value parameters of information data based on the value parameters of information data in multiple dimensions.

[0145] In some embodiments, the computer device determines the comprehensive value parameter of the information data based on the value parameters of the information data in multiple dimensions, including: the computer device weights the value parameters of the information data in multiple dimensions to obtain the comprehensive value parameter of the information data.

[0146] The weighting coefficients for the value parameters of different dimensions can be configured by the user or determined based on empirical values; this embodiment does not limit this. In some embodiments, the sum of the weighting coefficients for the value parameters of multiple dimensions is 1.

[0147] For example, the comprehensive value parameter of information data = H1 + H2 + H3 + H4 + H5 + H6. Where Hi is the value parameter of the information data in the i-th dimension, i = 1, 2, 3, 4, 5, 6. Hi = h1*w1 + h2*w2 + h3*w3, where h1 is the first evaluation indicator, w1 is the weight of the first evaluation indicator, h2 is the second evaluation indicator, w2 is the weight of the second evaluation indicator, and so on.

[0148] 305. Computer equipment displays the comprehensive value parameters of information data based on the information data.

[0149] In some embodiments, the computer device displays the information data based on a comprehensive value parameter, including: displaying the multiple pieces of information data in descending order of their value parameters. In some embodiments, the computer device displays the information data based on a comprehensive value parameter, including: displaying the information data if its value parameter is not less than a second threshold. In some embodiments, the computer device displays the information data based on a comprehensive value parameter, including: the computer device determines information data with value parameters not less than the second threshold, and displays the multiple pieces of information data in descending order of their determined value parameters.

[0150] It should be noted that the embodiments of this application are merely illustrative examples of how a computer device displays information data based on the value parameters of the information data. In other embodiments, the computer device may also price the information data or recommend the information data based on the value parameters. The embodiments of this application do not limit the application scenarios of the value parameters.

[0151] It should be noted that the embodiments in this application are merely illustrative examples of displaying information data based on comprehensive value parameters. In another embodiment, for each dimension of the information data's value parameters, a list of information data corresponding to that dimension can be displayed, allowing users to selectively obtain high-value information data across different dimensions.

[0152] The information data processing method provided in this application determines the value parameters of the information data after collecting it. These value parameters are determined based on the number of hot words belonging to the field and the frequency of their occurrence, ensuring the accuracy of the determined value parameters. By displaying the information data based on these value parameters, high-value information data can be presented to users, improving the accuracy of the information data presented to them.

[0153] Among them, the value parameters of information data are multi-dimensional. By evaluating the value of information data from multiple dimensions such as domain relevance, information credibility, content maturity, information dissemination, information attention, and information search popularity, the comprehensive value of information data can be accurately obtained, thus improving the accuracy of the value parameters.

[0154] Figure 4 This is a schematic diagram of the structure of an information data processing device provided in an embodiment of this application, such as... Figure 4 As shown, the device includes:

[0155] The data acquisition module 401 is used to collect information data belonging to the first domain.

[0156] The first determining module 402 is used to determine the number of hot words and the number of times hot words appear in the information data based on the reference data corresponding to the first domain. The reference data includes multiple hot words corresponding to the first domain. The number of hot words in the information data is the number of hot words in the information data that exist among the multiple hot words corresponding to the first domain. The number of times hot words appear is the number of times hot words appear in the information data.

[0157] The second determining module 403 is used to determine the value parameters of the information data based on the number of hot words and the frequency of hot word occurrences in the information data.

[0158] The display module 404 is used to display the information data based on the value parameters of the information data.

[0159] like Figure 5 As shown, in one possible implementation, the second determining module 403 includes:

[0160] Keyword extraction unit 4031 is used to extract keywords from the information data to obtain the keywords of the information data;

[0161] The value parameter determination unit 4032 is used to determine the value parameters of the information data based on the number of hot words, the number of times the hot words appear, and the number of times the keywords appear in the information data.

[0162] In one possible implementation, the second determining module 403 is used to determine the value parameters of the information data in multiple dimensions, including domain relevance, which is determined based on the number of hot words and the frequency of occurrence of hot words in the information data.

[0163] In one possible implementation, the value parameters of the multiple dimensions also include at least one of information credibility, content maturity, information dissemination, information attention, or information search popularity.

[0164] In one possible implementation, the credibility of the information is determined based on at least one of the data source of the information data and the number of characters in the body of the information data; or...

[0165] The content maturity level is determined based on at least one of the following: the frequency of occurrence of keywords in the information data and the relevance between the semantics of the title and the semantics of the body text of the information data; or...

[0166] The dissemination of the information is determined based on at least one of the following: the number of media outlets publishing the information data, the number of times the information data is shared, and the number of times the information data is forwarded; or...

[0167] The level of attention given to the information is determined based on at least one of the number of times the information data is saved and the number of times it is liked; or...

[0168] The search popularity of the information is determined based on the search information corresponding to the keywords in the information data.

[0169] In one possible implementation, the reference data further includes multiple data sources and the credibility of the multiple data sources; the second determining module 403 includes:

[0170] A value parameter determination unit is used to obtain a first credibility level corresponding to the data source of the information data from the reference data; determine a second credibility level of the information data based on the number of characters in the main body of the information data, wherein the second credibility level is positively correlated with the number of characters in the main body; and determine the credibility level of the information based on at least one of the first credibility level and the second credibility level.

[0171] In one possible implementation, the device further includes:

[0172] The acquisition module 401 is also used to acquire multiple pieces of information data belonging to the first domain;

[0173] The keyword extraction module 405 is used to extract keywords from the multiple pieces of information data to obtain the keywords of the multiple pieces of information data;

[0174] The hot word determination module 406 is used to use the keywords of the multiple pieces of information data as hot words corresponding to the first domain; or, to use the keywords of the multiple pieces of information data that appear at least once a first threshold as hot words corresponding to the first domain.

[0175] In one possible implementation, the device further includes:

[0176] The display module 404 is used to display a data configuration interface, which is used to obtain reference data corresponding to the first domain.

[0177] The reference data acquisition module 407 is used to use the data input in the data configuration interface as the reference data corresponding to the first field.

[0178] In one possible implementation, the display module 404 is used to display the multiple pieces of information data in descending order of their value parameters; or,

[0179] The display module 404 is used to display the information data when the value parameter of the information data is not less than a second threshold.

[0180] It should be noted that the information data processing device provided in the above embodiments is only illustrated by the division of the above functional modules when processing information data. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the information data processing device and the information data processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0181] Figure 6 This is a structural block diagram of an electronic device 600 provided in an embodiment of this application. The electronic device 600 includes a processor 601 and a memory 602.

[0182] Processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0183] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 are used to store at least one program code, which is executed by the processor 601 to implement the information data processing method provided in the method embodiments of this application.

[0184] In some embodiments, the electronic device 600 may optionally include a peripheral device interface 603 and at least one peripheral device. The processor 601, memory 602, and peripheral device interface 603 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 603 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 604, a display screen 605, a camera 606, an audio circuit 607, a positioning component 608, and a power supply 609.

[0185] Peripheral interface 603 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 601 and memory 602. In some embodiments, processor 601, memory 602 and peripheral interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 601, memory 602 and peripheral interface 603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0186] Display screen 605 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 605 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 601 for processing. In this case, display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, which serves as the front panel of electronic device 600; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of electronic device 600 or in a folded design; in still other embodiments, display screen 605 may be a flexible display screen, disposed on a curved or folded surface of electronic device 600. Furthermore, display screen 605 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 605 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0187] Power supply 609 is used to supply power to various components in electronic device 600. Power supply 609 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 609 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0188] Those skilled in the art will understand that Figure 6 The structure shown does not constitute a limitation on the electronic device 600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0189] Figure 7 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 700 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 701 and one or more memories 702. The memory 702 stores at least one line of program code, which is loaded and executed by the processor 701 to implement the methods provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.

[0190] The server 700 is used to execute the steps performed by the server in the above method embodiments.

[0191] This application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the information data processing method as described in any of the above implementations.

[0192] This application also provides a computer program product, which includes at least one piece of program code that is loaded and executed by a processor to implement the information data processing method as described in any of the above implementations.

[0193] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.

[0194] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An information data processing method, characterized in that, The method includes: Collect information data belonging to the first domain; Based on the reference data corresponding to the first domain, the value parameters of the information data in multiple dimensions are determined. The value parameters in multiple dimensions include domain relevance, information credibility, content maturity, information dissemination, information attention, and information search popularity. The reference data includes multiple data sources, the credibility of the multiple data sources, and multiple hot words corresponding to the first domain. The information search popularity is determined based on the search information corresponding to the keywords of the information data. The information data is weighted according to the value parameters of the multiple dimensions to obtain the comprehensive value parameters of the information data. The information data is displayed based on the comprehensive value parameters of the information data. The process of determining the domain relevance of the information data based on reference data corresponding to the first domain includes: Based on multiple hot words corresponding to the first domain, the number of hot words and the frequency of hot word occurrences in the information data are determined. The number of hot words in the information data is the number of hot words existing in the information data among the multiple hot words corresponding to the first domain, and the frequency of hot word occurrences is the number of times the hot words appear in the information data. Based on multiple domain words corresponding to the first domain, the number of domain words and the frequency of domain words occurrences in the information data are determined. The number of hot words, the frequency of hot word occurrences, the number of domain words, and the frequency of domain word occurrences are normalized to obtain a hot word count index, a hot word frequency index, a domain word count index, and a domain word frequency index, respectively. The hot word count index, the hot word frequency index, the domain word count index, and the domain word frequency index are weighted to obtain the domain relevance of the information data. The process of determining the information credibility of the information data based on the reference data corresponding to the first domain includes: The first credibility of the information data is obtained from the reference data; the second credibility of the information data is determined based on the number of characters in the main body of the information data, and the second credibility is positively correlated with the number of characters in the main body; the first credibility and the second credibility are weighted to obtain the information credibility.

2. The method according to claim 1, characterized in that, The method further includes: Collect multiple pieces of information data belonging to the first domain; Keyword extraction is performed on the multiple pieces of information data to obtain the keywords of the multiple pieces of information data; The keywords of the multiple pieces of information data are used as hot words corresponding to the first domain; or, the keywords of the multiple pieces of information data that appear at least once a first threshold number of times are used as hot words corresponding to the first domain.

3. The method according to any one of claims 1 to 2, characterized in that, The method further includes: Display a data configuration interface, which is used to obtain reference data corresponding to the first domain; The data entered in the data configuration interface is used as the reference data corresponding to the first domain.

4. The method according to any one of claims 1 to 2, characterized in that, The comprehensive value parameters based on the information data, which display the information data, include: The multiple pieces of information data are displayed in descending order of their comprehensive value parameters; or, The information data is displayed if the comprehensive value parameter of the information data is not less than the second threshold.

5. An information data processing device, characterized in that, The device includes: The data acquisition module is used to collect information data belonging to the first domain. The first determining module is used to determine the value parameters of the information data in multiple dimensions based on reference data corresponding to the first domain. These multiple value parameters include domain relevance, information credibility, content maturity, information dissemination, information attention, and information search popularity. The reference data includes multiple data sources, the credibility of the multiple data sources, and multiple hot words corresponding to the first domain. The information search popularity is determined based on search information corresponding to the keywords in the information data. Furthermore, the first determining module is used to determine the number of hot words and the frequency of hot word occurrences in the information data based on the multiple hot words corresponding to the first domain. The number of hot words in the information data is the number of hot words present in the information data among the multiple hot words corresponding to the first domain, and the frequency of hot word occurrences is... The frequency of occurrence of hot words in the information data is now determined; based on multiple domain words corresponding to the first domain, the number of domain words and the frequency of occurrence of domain words in the information data are determined. The number of domain words in the information data is the number of domain words existing in the information data among the multiple domain words corresponding to the first domain, and the frequency of occurrence of domain words is the number of times the domain words appear in the information data. The number of hot words, the frequency of occurrence of hot words, the number of domain words, and the frequency of occurrence of domain words are normalized to obtain the hot word count index, the hot word frequency index, the domain word count index, and the domain word frequency index, respectively. The hot word count index, the hot word frequency index, the domain word count index, and the domain word frequency index are weighted to obtain the domain relevance of the information data. The first determining module is used to obtain a first credibility level corresponding to the data source of the information data from the reference data; determine a second credibility level of the information data based on the number of characters in the main body of the information data, wherein the second credibility level is positively correlated with the number of characters in the main body; and perform weighted processing on the first credibility level and the second credibility level to obtain the information credibility level. The second determining module is used to weight the value parameters of the information data in multiple dimensions to obtain the comprehensive value parameters of the information data. The display module is used to display the information data based on the comprehensive value parameters of the information data.

6. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one piece of program code, which is loaded and executed by the processor to implement the information data processing method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the information data processing method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Information processing method and apparatus

    CN106933993A

  • Financial information acquisition method

    CN114818664A