Information identification methods, devices, electronic equipment and media
Patent Information
- Application Number
- CN202310181972.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-02-27
AI Technical Summary
[0005]本申请提供一种资讯信息的识别方法、装置、电子设备及介质,用以解决现有技术识别的资讯信息容易出现错误甚至丢失,导致识别的结果不够准确的问题
[0049] The information identification method, apparatus, electronic device, and medium provided in this application acquire information from a displayed page, perform segmentation identification processing on the information, obtain source-related information, and perform structured processing on the source-related information to obtain source information with a preset structure. This preset structure includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time. Then, based on a preset publisher database, platform account database, platform database, and author database, the source-related content in the preset structure of source information is identified to determine the target source of the information. The target source includes at least one of the publisher, original source, and author. The method of this application improves the accuracy of the identification results by performing segmentation identification processing, structured processing, and matching identification with a preset database on the information.
Smart Images

Figure CN116186405B_ABST
Abstract
Description
Technical Field
[0001] This application relates to text recognition technology, and more particularly to a method, apparatus, electronic device, and medium for recognizing information. Background Technology
[0002] With the rapid development of the information age, more and more information is appearing in front of users, and users can obtain useful information in a relatively short time by reading the information.
[0003] In existing technologies, when a terminal identifies information on a displayed page, it typically simply identifies the information as is and then stores the identified information in a background table.
[0004] However, existing technologies are prone to errors or even loss of information, resulting in inaccurate identification results and an inability to effectively distinguish between information such as publishers, original sources, and authors that are mixed together. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and medium for identifying information, in order to solve the problem that the information identified by the prior art is prone to errors or even loss, resulting in inaccurate identification results.
[0006] Firstly, this application provides a method for identifying information, including:
[0007] Retrieve information from the displayed page;
[0008] The information is partitioned and identified to obtain source-related information.
[0009] The source-related information is structured to obtain source information with a preset structure, which includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time.
[0010] Based on a preset publisher database, platform account database, platform database, and author database, source-related content in the source information of the preset structure is identified to determine the target source of the information. The target source includes at least one of the publisher, the original source, and the author.
[0011] Optionally, the step of performing partition identification processing on the information to obtain source-related information of the information includes:
[0012] The display page is divided into a preset number of display areas, and the information displayed in each display area is determined.
[0013] Based on the information displayed in each display area, the information is divided into attributes to determine the attribute area where each display area is located. The attribute area includes title area, account area, publication archive area, summary area, body text area, body text notes area, and comment area.
[0014] Identify the source-related information of the information from the attribute area.
[0015] Optionally, if the target source is the publisher;
[0016] The method, based on a preset publisher database, platform account database, platform database, and author database, identifies source-related content in the source information within the preset structure to determine the target source of the information, including:
[0017] Obtain the publication link of the information in the source information of the preset structure;
[0018] The publishing information in the publishing link is matched with the publishing platforms stored in the preset platform database. If the match is consistent, the information is published by the platform; if the match is inconsistent, the information is published by a non-platform.
[0019] Obtain the publishing platform account information from the source information of the preset structure;
[0020] If the information is published by the platform, the platform account information will be matched with the platform account information stored in the preset platform account database;
[0021] If a match is found, the publisher corresponding to the publishing platform account information is retrieved from the preset platform account database, and the publisher is matched with the publishers stored in the preset publisher database. If a match is found, the target source is the stored publisher.
[0022] If the information is not published by the platform, obtain the publisher from the source information in the preset structure;
[0023] The publisher is matched with publishers stored in a preset publisher database. If a match is found, the target source is the stored publisher.
[0024] Optionally, if the target source is the original source;
[0025] The method, based on a preset publisher database, platform account database, platform database, and author database, identifies source-related content in the source information within the preset structure to determine the target source of the information, including:
[0026] Obtain the original source from the source information of the preset structure;
[0027] The original source is matched with the publisher information stored in the preset publisher database;
[0028] If a match is found, the target source is the original source in the obtained source information.
[0029] Optionally, if the target source is the author;
[0030] The method, based on a preset publisher database, platform account database, platform database, and author database, identifies source-related content in the source information within the preset structure to determine the target source of the information, including:
[0031] Obtain the author from the source information of the preset structure;
[0032] The authors are matched with authors stored in a preset author database;
[0033] If a match is found, the target source is the author in the obtained source information;
[0034] If the match is inconsistent, semantic analysis is performed on the authors in the obtained source information to obtain an analysis score.
[0035] If the analysis score is greater than a preset score threshold, then the target source is the author in the obtained source information.
[0036] Optional, also includes:
[0037] The target source of the determined information is output, and the target source includes at least one of the publisher, the original source, and the author.
[0038] Secondly, this application provides an information identification device, comprising:
[0039] The acquisition module is used to retrieve information from the displayed page.
[0040] The identification module is used to perform partition identification processing on the information and obtain source-related information of the information;
[0041] The processing module is used to perform structured processing on the source-related information to obtain source information with a preset structure. The preset structure includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time.
[0042] The identification module is also used to identify source-related content in the source information of the preset structure based on a preset publisher database, platform account database, platform database, and author database, and to determine the target source of the information, wherein the target source includes at least one of publisher, original source, and author.
[0043] Thirdly, this application provides an electronic device, comprising:
[0044] At least one processor and memory;
[0045] The memory stores computer-executed instructions;
[0046] The at least one processor executes computer execution instructions stored in the memory, causing the electronic device to perform the information identification method according to any one of the first aspects.
[0047] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the information identification method as described in any of the first aspects.
[0048] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the information identification method described in any of the first aspects.
[0049] The information identification method, apparatus, electronic device, and medium provided in this application acquire information from a displayed page, perform segmentation identification processing on the information, obtain source-related information, and perform structured processing on the source-related information to obtain source information with a preset structure. This preset structure includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time. Then, based on a preset publisher database, platform account database, platform database, and author database, the source-related content in the preset structure of source information is identified to determine the target source of the information. The target source includes at least one of the publisher, original source, and author. The method of this application improves the accuracy of the identification results by performing segmentation identification processing, structured processing, and matching identification with a preset database on the information. Attached Figure Description
[0050] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0051] Figure 1A flowchart illustrating an information identification method provided in an embodiment of this application;
[0052] Figure 2 A flowchart illustrating a method for partitioning and identifying information provided in this application embodiment;
[0053] Figure 3 A flowchart illustrating a method for determining the publisher of information provided in this application embodiment;
[0054] Figure 4 A flowchart illustrating a method for determining the original source of information provided in an embodiment of this application;
[0055] Figure 5 A flowchart illustrating a method for determining the author of information provided in an embodiment of this application;
[0056] Figure 6 An information identification device provided in this application embodiment;
[0057] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0058] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0059] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0060] In the description of the embodiments in this application, terms such as "first," "second," and "third" are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can also be implemented in sequences other than those illustrated or described in this application. Terms such as "inner" and "outer" that indicate directional or positional relationships are based on the directional or positional relationships shown in the drawings and are merely for ease of description, and do not indicate or imply that a device or component must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as limiting this application.
[0061] Furthermore, in the description of the embodiments of this application, unless otherwise expressly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the embodiments of this application according to the specific circumstances.
[0062] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0063] First, let me explain the terms used in this application:
[0064] Publisher: This refers to the entity that publishes information, such as traditional media, self-media, or individuals. Other synonyms for publisher in practical applications include, but are not limited to: publishing media, publishing media, affiliated media, publisher, publishing entity, publisher, etc.
[0065] Traditional media refers to media presented in traditional forms, often appearing in an organized manner, including but not limited to newspapers, magazines, television, radio, and online media.
[0066] Self-media refers to media formed by organizations or individuals registering accounts on publishing platforms after the rise of social networks, including but not limited to WeChat official accounts, Toutiao accounts, etc.
[0067] Original source: refers to the original publisher of information. For example, if website A publishes an article from newspaper B, then newspaper B is the original source. Other synonyms for original source in practical applications include, but are not limited to: source media, reprinted from, source party, original publishing media, original publisher, etc.
[0068] Author: refers to a natural person who writes or provides information content. Other synonyms for author in practical application include, but are not limited to: writer, author, text / content provider, etc.
[0069] Publishing platform: refers to the type of platform on which information is published. Other synonyms for publishing platform in practical applications include, but are not limited to: platform media, public account, self-media platform, etc.
[0070] Publishing platform account: refers to the account name registered by a user on the platform. Other synonyms for publishing platform account in actual application include, but are not limited to: registered account, account name, nickname, platform account, registered name, etc.
[0071] Channel: refers to the media through which a piece of information is ultimately published. For example, if a piece of information is published on platform A and then automatically appears on media B, which is linked from platform A, then media B is the channel. Other synonyms for channel in practical applications include, but are not limited to: channel media, channel, exposure media, etc.
[0072] The rapid development of the information age has led to an increasing amount of information appearing in front of users. Users obtain the knowledge they need by browsing the information displayed on the terminal display page, and identify the content they are interested in by identifying the displayed information.
[0073] Currently, when terminals identify information on displayed pages, they typically simply recognize the displayed information as is and then store the identified information in a background table. However, the current identification method is not accurate enough.
[0074] For example, current identification methods are prone to data loss when identifying the publisher, original source, and author of information. Furthermore, during the identification process, they may mistakenly identify author information as publisher information or vice versa, leading to identification errors. Moreover, current methods can only recognize the content of the displayed information as is, and cannot perform further in-depth identification based on the displayed information. For instance, because source-related information is mixed in, current methods can identify where the information was published, but cannot identify the publisher, resulting in incomplete identification.
[0075] Therefore, to address the aforementioned technical problems of the prior art, this application proposes a method, apparatus, electronic device, and medium for identifying information. By performing partitioning and identification processing on the information, source-related information is obtained. The obtained source-related information is then structured to obtain source information with a preset structure. Based on multiple preset databases, the source-related content within the preset structured source information is identified, ultimately determining the target source of the information. The target source includes at least one of a publisher, original source, and author. This method improves the accuracy of information identification.
[0076] The application scenario of this application can be the identification of information displayed on a terminal, wherein the information can be textual materials such as news, technology, strategy, commentary, opinions, and academic data. It is understood that the information may also include other types of text; the above examples are for illustrative purposes only and are not intended to limit this application.
[0077] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0078] Figure 1 This is a flowchart illustrating an information identification method provided in an embodiment of this application. The execution subject of this method can be a terminal with information identification capabilities, such as a personal computer, laptop, smartphone, tablet, server, and portable wearable device. The method in this embodiment can be implemented through software, hardware, or a combination of both. Figure 1 As shown, the method specifically includes the following steps:
[0079] S101. Obtain information from the displayed page.
[0080] In this embodiment, the executing entity is a computer. The computer obtains information from the currently displayed page. The information can be plain text or a combination of text and images, such as an article or a report.
[0081] S102. Perform partition identification processing on the information to obtain relevant information about the source of the information.
[0082] The computer partitions the information into sections and retrieves the source-related information for each section based on the information in that section.
[0083] The divided areas include, but are not limited to: title area, account area, posting archive area, summary area, body text area, body text notes area, and comment area. It is understood that this embodiment does not limit the specific areas to be divided.
[0084] Source-related information can include information related to the publisher of the information, information related to the original source of the information, information related to the author of the information, and so on.
[0085] Partitioning method:
[0086] Optionally, the section can be determined based on the identifiers in the information. For example, the section can be determined as the "Account" section based on the WeChat Official Account identifier in the information.
[0087] Optionally, the sections can be determined based on the text format of the information. For example, the sentences with the largest font size and no period in the information can be designated as the title section.
[0088] Optionally, sections can be defined based on keywords in the information. For example, paragraphs containing the keyword "abstract" can be designated as the abstract section.
[0089] The above partitioning method is for illustrative purposes only and is not intended to limit this application.
[0090] S103. Perform structured processing on the source-related information to obtain source information with a preset structure.
[0091] The preset structure includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time.
[0092] To facilitate the identification of information sources, it is necessary to perform structured processing.
[0093] For example,
[0094] If the source information is related to where the news information was published, then the corresponding content will be categorized into a field, such as "Local Media".
[0095] If the source information is "from which source" or "original source", the content corresponding to "from which source" or "original source" will be divided into a field, such as "Source Media".
[0096] If the source information includes both the publisher and the writer, the content related to the publisher should be divided into one field, such as "Publisher", and the content related to the writer should be divided into another field, such as "Author".
[0097] S104. Based on the preset publisher database, platform account database, platform database, and author database, identify source-related content in the source information of the preset structure, and determine the target source of the information. The target source includes at least one of the publisher, the original source, and the author.
[0098] After obtaining the source information of the preset structure in step S103, the source-related content in the source information is matched and identified with the preset database to determine at least one of the publisher, original source and author.
[0099] In the above embodiments of this application, by acquiring information information from the displayed page and performing segmentation and identification processing on the information information, source-related information of the information information is obtained. This source-related information is then structured to obtain source information with a preset structure. The preset structure includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time. Finally, based on preset publisher databases, platform account databases, platform databases, and author databases, the source-related content in the source information of the preset structure is identified to determine the target source of the information. The target source includes at least one of the publisher, original source, and author. The method of this embodiment improves the accuracy of information information identification results.
[0100] Furthermore, based on the above embodiments, the process of performing partition identification processing on information information and obtaining source-related information of information information in step S102 is explained through the following embodiments. Figure 2 This application provides a flowchart illustrating a method for partitioning and identifying information, as shown in the embodiments below. Figure 2 As shown, the method includes the following steps:
[0101] S201. Divide the display page into a preset number of display areas and determine the information to be displayed in each display area.
[0102] Optionally, the display page can be divided into a preset number of display areas, or randomly divided into a preset number of display areas. The preset number can be determined based on the amount of information, or set by the user according to the actual situation.
[0103] After dividing the display area into a preset number, determine the information displayed in each display area.
[0104] S202. Based on the information displayed in each display area, classify the information by attributes and determine the attribute area where each display area is located.
[0105] The attribute area includes, but is not limited to: title area, account area, publication archive area, summary area, body area, body notes area, and comment area.
[0106] S203. Identify the source-related information of the information from the attribute area.
[0107] For example, you can obtain the publishing platform account information in the account area.
[0108] For example, you can retrieve the source of information, author, publisher, and publication time of information from the published archive area.
[0109] For example, you can retrieve preset keyword information from the notes section of the main text, and determine the source and author of the information based on the preset keyword information.
[0110] It should be noted that the source-related information can only be obtained if the attribute area exists; different attribute areas may yield the same source-related information.
[0111] In the above embodiments of this application, the display page is divided into a preset number of display areas, the information displayed in each display area is determined, and the information is attribute-divided according to the information displayed in each display area to determine the attribute area where each display area is located. Then, the source-related information of the information is identified from the attribute area. This embodiment, by performing partitioning and identification processing on the information, makes the identified source-related information more comprehensive and complete.
[0112] Furthermore, the following embodiments illustrate the process in step S104 of identifying source-related content in the source information of the preset structure based on the preset publisher database, platform database, and author database, and determining the target source of the information.
[0113] If the target source is the publisher Figure 3 A flowchart illustrating a method for determining the publisher of information provided in this application embodiment is shown below. Figure 3 As shown, the method includes the following steps:
[0114] S301. Obtain the publishing link of information information in the source information of the preset structure.
[0115] S302. Match the publishing information in the publishing link with the publishing platforms stored in the preset platform database. If the match is consistent, the information is published by the platform. If the match is inconsistent, the information is published by a non-platform.
[0116] The system pre-stores different types of platforms in a pre-defined platform database, such as a WeChat official account or a Toutiao account.
[0117] The publishing link includes publishing information, which can be, for example, the identification information of a certain public account. This identification information is matched with the platforms in the preset platform database. If they match, the information is published by the platform; otherwise, it is published by a non-platform.
[0118] Alternatively, a pre-defined channel database can be used to determine the platform for publishing information. Here, a channel refers to the media where a piece of information is ultimately published. For example, if a piece of information is published on platform A and ultimately displayed on platform B, which is related to a certain platform, then platform B is the channel.
[0119] S303. Obtain the publishing platform account information from the source information of the preset structure.
[0120] The platform account information refers to the user's registered account name, such as nickname.
[0121] S304. If the information is published by the platform, the platform account information will be matched with the platform account information stored in the preset platform account database. If the match is found, the publisher corresponding to the platform account information will be retrieved from the preset platform account database. The publisher will be matched with the publisher stored in the preset publisher database. If the match is found, the target source will be the stored publisher.
[0122] The system pre-stores user registration account information in a pre-defined platform account database. Each account has a corresponding publisher. Therefore, after obtaining the publishing platform account information from the source information with the pre-defined structure, the publisher corresponding to it can be determined through the platform account database.
[0123] Correspondingly, different publishers are pre-stored in the preset publisher database. Therefore, if the publisher identified in the previous step can be matched in the publisher database, the target source is the stored publisher.
[0124] S305. If the information is published by a non-platform, obtain the publisher from the source information in the preset structure.
[0125] S306. Match the publisher with the publishers stored in the preset publisher database. If they match, the target source is the stored publisher.
[0126] pass Figure 3 The method shown effectively improves the accuracy of publisher identification in consultation information.
[0127] If the target source is the original source. Figure 4 A flowchart illustrating a method for determining the original source of information provided in this application embodiment is shown below. Figure 4 As shown, the method includes the following steps:
[0128] S401. Obtain the original source from the source information of the preset structure.
[0129] S402. Match the original source with the publisher information stored in the preset publisher database.
[0130] The publisher data also stores the original sources of different information in advance. Each original source has its corresponding publisher. Therefore, once the publisher is determined, the original source corresponding to that publisher in the publisher database can be used to determine whether the original source in the source information of the preset structure is the real original source of the information.
[0131] S403. If a match is found, the target source is the original source in the obtained source information.
[0132] The original source in the source information of the obtained preset structure is matched with the publisher information in the publisher database. If a match is found, it means that the target source is the original source in the obtained source information.
[0133] pass Figure 4 The method shown effectively improves the accuracy of identifying the original source in consultation information.
[0134] If the target source is the author. Figure 5 A flowchart illustrating a method for determining the author of information provided in this application embodiment is shown below. Figure 5 As shown, the method includes the following steps:
[0135] S501. Obtain the author from the source information of the preset structure.
[0136] S502. Match the author with the authors stored in the preset author database.
[0137] S503. If a match is found, the target source is the author in the obtained source information.
[0138] The author database stores the authors. If a matching author can be found in the author database, it means that the target source is the author in the source information obtained.
[0139] For example, assuming the author in the source information of the obtained preset structure is Aaa, and the author Aaa is stored in the author database, then the target source is author Aaa.
[0140] It should be noted that if the author database includes multiple Aaa authors, you can confirm their identity through other information related to Aaa in the author database, such as the author Aaa's place of origin and works.
[0141] S504. If the match is inconsistent, semantic analysis is performed on the authors in the obtained source information to obtain the analysis score.
[0142] If the matching is inconsistent, semantic analysis can be performed on the authors in the obtained source information using a preset algorithm or model to obtain the corresponding analysis score.
[0143] S505. If the analysis score is greater than the preset score threshold, the target source is the author in the obtained source information.
[0144] For example, suppose the author in the source information of the preset structure is Bbbb. This author does not appear in the author database, but through semantic analysis, the analysis score obtained is 90%, which is greater than the preset score threshold of 80%. This indicates that Bbbb is very likely to be a natural person's name. Therefore, Bbbb is identified as the target source, that is, the target source is the author in the source information.
[0145] pass Figure 5 The method shown effectively improves the accuracy of author identification in consultation information.
[0146] In summary, after determining the target source of the information using the above methods, the determined target source is output, which includes at least one of the publisher, original source, and author.
[0147] If the above method fails to identify the corresponding publisher, original source, or author based on the preset database, the output will be abnormal data.
[0148] In the above embodiments of this application, source-related content in source information with a preset structure is identified by using a preset publisher database, platform account database, platform database, author database, etc., so that the publisher, original source or author in the identified information is more accurate.
[0149] Figure 6 An information identification device provided in this application embodiment, such as Figure 6 As shown, the device includes: an acquisition module 601, an identification module 602, and a processing module 603.
[0150] The acquisition module 601 is used to acquire information from the display page.
[0151] The identification module 602 is used to perform partition identification processing on the information and obtain the source information of the information.
[0152] The processing module 603 is used to perform structured processing on the source-related information to obtain source information with a preset structure. The preset structure includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time.
[0153] The identification module 602 is also used to identify source-related content in the source information of the preset structure based on the preset publisher database, platform account database, platform database, and author database, and to determine the target source of the information. The target source includes at least one of the publisher, the original source, and the author.
[0154] One possible implementation is that the identification module 602 is specifically used for:
[0155] The page will be divided into a preset number of display areas, and the information displayed in each display area will be determined.
[0156] Based on the information displayed in each display area, the information is divided into attributes to determine the attribute area where each display area is located. The attribute areas include title area, account area, publication archive area, summary area, body text area, body text notes area, and comment area.
[0157] Identify the source information of the information from the attribute area.
[0158] In one possible implementation, the identification module 602 is further used for:
[0159] Retrieve the publication links of information from the source information of the preset structure.
[0160] The information in the publishing link is matched with the publishing platforms stored in the preset platform database. If they match, the information is published by the platform; otherwise, it is published by a non-platform.
[0161] Retrieve the publishing platform account information from the source information of the preset structure.
[0162] If the information is published by the platform, the platform account information will be matched with the platform account information stored in the preset platform account database.
[0163] If a match is found, the publisher corresponding to the publishing platform account information is retrieved from the preset platform account database. The publisher is then matched with the publishers stored in the preset publisher database. If a match is found, the target source is the stored publisher.
[0164] If the information is published by a non-platform entity, retrieve the publisher from the source information in the preset structure.
[0165] The publisher is matched with the publishers stored in the preset publisher database. If a match is found, the target source is the stored publisher.
[0166] In one possible implementation, the identification module 602 is further used for:
[0167] Obtain the original source from the source information of the preset structure.
[0168] Match the original source with the publisher information stored in the preset publisher database.
[0169] If a match is found, the target source is the original source in the obtained source information.
[0170] In one possible implementation, the identification module 602 is further used for:
[0171] Retrieve the author from the source information of the preset structure.
[0172] The author is matched with authors stored in a pre-defined author database.
[0173] If a match is found, the target source is the author in the obtained source information.
[0174] If the match is inconsistent, semantic analysis is performed on the authors in the obtained source information to obtain the analysis score.
[0175] If the analysis score is greater than the preset score threshold, the target source is the author in the obtained source information.
[0176] In one possible implementation, the identification module 602 is also used for:
[0177] The target source of the output information is determined, and the target source includes at least one of the publisher, original source, and author.
[0178] The information recognition device provided in this embodiment is used to execute the aforementioned method embodiment. Its implementation principle and technical effect are similar, and will not be described again.
[0179] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 7 As shown, the device may include at least one processor 701 and a memory 702.
[0180] The memory 702 is used to store programs. Specifically, the program may include program code, which may include computer operation instructions or executable instructions of the processor 701, etc.
[0181] The memory 702 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0182] The processor 701 is used to execute computer execution instructions stored in the memory 702 to implement the content described in the foregoing method embodiments. The processor 701 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0183] Optionally, the electronic device may also include a communication interface 703. In specific implementations, if the communication interface 703, memory 702, and processor 701 are implemented independently, they can be interconnected via a bus to complete communication. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc., but this does not imply that there is only one bus or one type of bus.
[0184] Optionally, in a specific implementation, if the communication interface 703, memory 702, and processor 701 are integrated on a single chip, then the communication interface 703, memory 702, and processor 701 can communicate through an internal interface.
[0185] The electronic device provided in this embodiment is used to execute the information recognition method executed in the aforementioned embodiment. Its implementation principle and technical effect are similar, and will not be described again.
[0186] This application also provides a computer-readable storage medium, which may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a disk, or an optical disk. Specifically, the computer-readable storage medium stores computer-executable instructions, which are used for the information identification method in the above embodiments.
[0187] This application also provides a computer program product comprising executable instructions or a computer program stored in a readable storage medium. At least one processor of an electronic device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to implement the aforementioned information identification method.
[0188] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0189] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for identifying information, characterized in that, include: Obtain information about the user's browser terminal's displayed page; The display page is divided into a preset number of display areas, and the information displayed in each display area is determined. Based on the information displayed in each display area, the information is divided into attributes to determine the attribute area where each display area is located. The attribute area includes a title area, account area, publication archive area, summary area, body text area, body text notes area, and comment area. The source-related information of the information is identified from the attribute area. The source-related information includes information related to the original source of the information. The source-related information is structured to obtain source information with a preset structure, which includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time. Based on a preset publisher database, platform account database, platform database, and author database, the source-related content in the source information of the preset structure is identified to determine the target source of the information, including: obtaining the original source in the source information of the preset structure, matching the original source with the publisher information stored in the preset publisher database, and if the match is consistent, the target source is the original source in the obtained source information. The target source includes at least one of the publisher, the original source, and the author; The method, based on a preset publisher database, platform account database, platform database, and author database, identifies source-related content in the source information within the preset structure to determine the target source of the information, including: Obtain the publication link of the information in the source information of the preset structure; The publishing information in the publishing link is matched with the publishing platforms stored in the preset platform database. If the match is consistent, the information is published by the platform; if the match is inconsistent, the information is published by a non-platform. Obtain the publishing platform account information from the source information of the preset structure; If the information is published by the platform, the platform account information will be matched with the platform account information stored in the preset platform account database; If a match is found, the publisher corresponding to the publishing platform account information is retrieved from the preset platform account database, and the publisher is matched with the publishers stored in the preset publisher database. If a match is found, the target source is the stored publisher. If the information is not published by the platform, obtain the publisher from the source information in the preset structure; The publisher is matched with publishers stored in a preset publisher database. If a match is found, the target source is the stored publisher.
2. The method according to claim 1, characterized in that, If the source of the target is the author; The method, based on a preset publisher database, platform account database, platform database, and author database, identifies source-related content in the source information within the preset structure to determine the target source of the information, including: Obtain the author from the source information of the preset structure; The authors are matched with authors stored in a preset author database; If a match is found, the target source is the author in the obtained source information; If the match is inconsistent, semantic analysis is performed on the authors in the obtained source information to obtain an analysis score. If the analysis score is greater than a preset score threshold, then the target source is the author in the obtained source information.
3. The method according to claim 1 or 2, characterized in that, Also includes: The target source of the determined information is output, and the target source includes at least one of the publisher, the original source, and the author.
4. An information recognition device, characterized in that, include: The acquisition module is used to acquire information about the display page on the user's browsing terminal; The identification module is used to divide the display page into a preset number of display areas and determine the information displayed in each display area; Based on the information displayed in each display area, the information is divided into attributes to determine the attribute area where each display area is located. The attribute area includes a title area, account area, publication archive area, summary area, body text area, body text notes area, and comment area. The source-related information of the information is identified from the attribute area. The source-related information includes information related to the original source of the information. The processing module is used to perform structured processing on the source-related information to obtain source information with a preset structure. The preset structure includes information content, media, channel, author, publisher, original source, publishing platform, publishing platform account, publishing link, and publishing time. The identification module is further configured to identify source-related content in the source information of the preset structure based on a preset publisher database, platform account database, platform database, and author database, and determine the target source of the information, including: obtaining the original source in the source information of the preset structure, matching the original source with the publisher information stored in the preset publisher database, and if the match is consistent, the target source is the original source in the obtained source information; and obtaining the publishing link of the information in the source information of the preset structure. The process involves several steps: First, matching the publishing information in the publishing link with publishing platforms stored in a preset platform database. If a match is found, the information is published by the platform; otherwise, it is published by a non-platform. Then, the process retrieves the publishing platform account information from the source information within the preset structure. If the information is published by the platform, the platform account information is matched with platform account information stored in the preset platform account database. If a match is found, the publisher corresponding to the publishing platform account information is retrieved from the preset platform account database, and this publisher is matched with publishers stored in a preset publisher database. If a match is found, the target source is the stored publisher. If the information is published by a non-platform, the publisher is retrieved from the source information within the preset structure, and this publisher is matched with publishers stored in the preset publisher database. If a match is found, the target source is the stored publisher. The target source includes at least one of the publisher, the original source, and the author.
5. An electronic device, characterized in that, include: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the electronic device to perform the information identification method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the information identification method as described in any one of claims 1 to 3.
7. A computer program product, characterized in that, The system includes a computer program that, when executed by a processor, implements the information identification method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Automatic identification method for multimedia news original work and first media
CN112579800A
Real-time information crawler method and device based on intelligent page analysis, and equipment
CN113987320A