Bulletin title generation method, apparatus, terminal, and medium

By analyzing the text of stock announcement documents and selecting target line vectors to generate titles, the problem of stock announcement titles failing to accurately reflect the content is solved, thus improving the convenience for users to access announcement content.

CN115934892BActive Publication Date: 2026-02-17SHENZHEN FUTU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211584172.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2026-02-17
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

The titles of stock announcements fail to accurately reflect the content of the announcements, affecting the ease with which users can access the information.

Method used

By analyzing the text of the announcement document, target line vectors are selected based on document type, font type, and layout style to generate an accurate announcement title.

Benefits of technology

Ensure that the announcement title accurately reflects the announcement content, thereby improving the ease with which users can access the announcement information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115934892B_ABST
    Figure CN115934892B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of finance, and particularly relates to a kind of announcement title generation method, device, terminal and medium.The announcement title generation method comprises: determining to be analyzed file in the file of announcement according to file type;Text content of the to-be-analyzed file is converted into a set of row vectors;According to font type and layout style, the row vectors in the set of row vectors are filtered to determine target row vectors;Generate announcement title based on the target row vectors.In this way, the application analyzes the text content of the announcement file by text analysis, to obtain key content in the announcement file as the announcement title, so as to ensure that the announcement title can accurately reflect the announcement content, and improve the convenience of users to obtain announcement content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial technology, and in particular to a method, apparatus, terminal, and medium for generating announcement titles. Background Technology

[0002] The title is crucial for a piece of news. From the perspective of the news itself, the title is usually the most concise expression of the content; from the perspective of user experience, users can quickly understand the content of the news through the title while browsing.

[0003] However, in today's stock market, stock announcement titles often fail to accurately reflect the content of the corresponding stock announcement. For example, after US stock announcements are entered into the database, the titles of the corresponding financial statements or other official documents received by the US Securities and Exchange Commission (SEC) are usually used as the announcement titles. However, SEC document titles are announcement types and cannot accurately reflect the content of the announcement, thus affecting the convenience for users to access the announcement content.

[0004] Therefore, how to generate titles for stock announcements that accurately reflect the content of the announcements is a problem that urgently needs to be solved in the field of financial technology. Summary of the Invention

[0005] The main objective of this invention is to provide a method, apparatus, terminal, and medium for generating announcement titles. The aim is to analyze the text content of announcement documents to extract key information from the announcement documents and use it as the announcement title, thereby ensuring that the announcement title accurately reflects the announcement content.

[0006] According to one aspect of the embodiments of this application, a method for generating announcement titles is disclosed, including:

[0007] The file to be analyzed is determined from the files that make up the announcement based on the file type;

[0008] Convert the text content of the file to be analyzed into a set of line vectors;

[0009] The row vectors in the set of row vectors are filtered according to font type and typography style to determine the target row vectors;

[0010] The announcement title is generated based on the target row vector.

[0011] In some embodiments of this application, based on the above technical solutions, the text content of the file to be analyzed is converted into a set of line vectors, including:

[0012] Detect the tag information contained in the file to be analyzed;

[0013] Based on the tag information, the text content of the file to be analyzed is traversed and recursively converted into multiple line vectors;

[0014] A set of row vectors is generated based on the multiple row vectors.

[0015] In some embodiments of this application, based on the above technical solutions, the text content of the file to be analyzed is traversed recursively according to the tag information to convert the text content into multiple line vectors, including:

[0016] The main body of the text content of the file to be analyzed is determined based on the tag information.

[0017] The main body is traversed recursively based on preset segmentation labels and preset table labels to convert the text content into multiple row vectors.

[0018] In some embodiments of this application, based on the above technical solutions, the file to be analyzed is determined from the files constituting the announcement according to the file type, including:

[0019] The file types of the documents that make up the announcement are checked;

[0020] The target attachment file is determined from the files that make up the announcement based on the file type described;

[0021] The target attachment file is identified as the file to be analyzed.

[0022] In some embodiments of this application, based on the above technical solutions, the target supplementary file is determined from the files constituting the announcement according to the file type, including:

[0023] The additional documents are determined from the documents that make up the announcement, based on their document types;

[0024] The attached file is compared with a preset priority list to obtain the comparison result. The preset priority list is used to reflect the probability of obtaining the appropriate announcement title in various types of attached files.

[0025] The target attachment file is determined based on the comparison results.

[0026] In some embodiments of this application, based on the above technical solutions, the target supplementary file is determined from the files constituting the announcement according to the file type, including:

[0027] The main file is determined from the files that make up the announcement based on the file type;

[0028] Extract the target attachment information from the text content corresponding to the preset node in the main file;

[0029] The target attachment file is determined based on the target attachment information.

[0030] In some embodiments of this application, based on the above technical solutions, generating an announcement title based on the target row vector includes:

[0031] Convert the text content of the document that makes up the announcement into a preset set of words;

[0032] Identify target words within the preset word set that achieve a preset frequency of occurrence;

[0033] The target row vectors are prioritized according to the target words;

[0034] The announcement title is generated based on the target row vectors sorted by priority.

[0035] According to one aspect of the embodiments of this application, an announcement title generation device is disclosed, the announcement title generation device comprising:

[0036] The file determination module is configured to determine the file to be analyzed from the files that make up the announcement based on the file type;

[0037] The text conversion module is configured to convert the text content of the file to be analyzed into a set of line vectors;

[0038] The row vector filtering module is configured to filter row vectors in the row vector set according to font type and typography style to determine the target row vector;

[0039] The title generation module is configured to generate an announcement title based on the target row vector.

[0040] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the announcement header generation method as described in the above technical solutions.

[0041] The announcement title generation method provided in this application first determines the file to be analyzed from multiple files that make up the announcement based on the file type. The file to be analyzed is the one with a high probability of obtaining a suitable announcement title. Then, the text to be analyzed is split into a set of line vectors. Next, the text to be analyzed is analyzed by filtering the line vectors in the set according to the font type and layout style. The line vectors with a high probability of being suitable as announcement titles are selected as target vectors and used as candidate announcement titles. Finally, the actual announcement title is determined from the candidate announcement titles.

[0042] Thus, the announcement title generation method provided in this application analyzes the text content of the announcement document to obtain key content as the announcement title, thereby ensuring that the announcement title accurately reflects the announcement content and improving the convenience for users to obtain the announcement content.

[0043] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application. Attached Figure Description

[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0045] Figure 1 A flowchart illustrating the steps of a method for generating announcement titles according to one embodiment of this application is shown.

[0046] Figure 2 This illustration shows the effect of obtaining the target row vector in an appendix file in one embodiment of this application.

[0047] Figure 3 This illustration shows the effect of obtaining the target line vector from the text content corresponding to a specific node in the main file in one embodiment of this application.

[0048] Figure 4 A schematic block diagram of the announcement title generation apparatus provided in an embodiment of this application is shown.

[0049] Figure 5 A schematic diagram of a computer system architecture suitable for implementing the terminal device of the present application is shown. Detailed Implementation

[0050] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0051] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0052] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0053] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0054] The following detailed description of the technical solutions provided in this application, including the method, apparatus, terminal, and medium for generating announcement titles, is based on specific embodiments.

[0055] Figure 1 A flowchart illustrating the steps of an announcement title generation method in one embodiment of this application is shown, as follows: Figure 1 As shown, the method for generating the announcement title can mainly include the following steps S100 to S400.

[0056] Step S100: Determine the file to be analyzed from the files that make up the announcement based on the file type.

[0057] Step S200: Convert the text content of the file to be analyzed into a set of line vectors.

[0058] Step S300: Filter the row vectors in the set of row vectors according to font type and layout style to determine the target row vector.

[0059] Step S400: Generate an announcement title based on the target row vector.

[0060] The announcement title generation method provided in this application first determines the file to be analyzed from multiple files that make up the announcement based on the file type. The file to be analyzed is the one with a high probability of obtaining a suitable announcement title. Then, the text to be analyzed is split into a set of line vectors. Next, the text to be analyzed is analyzed by filtering the line vectors in the set according to the font type and layout style. The line vectors with a high probability of being suitable as announcement titles are selected as target vectors and used as candidate announcement titles. Finally, the actual announcement title is determined from the candidate announcement titles.

[0061] Thus, the announcement title generation method provided in this application analyzes the text content of the announcement document to obtain key content as the announcement title, thereby ensuring that the announcement title accurately reflects the announcement content and improving the convenience for users to obtain the announcement content.

[0062] The following sections provide a detailed explanation of each step in the method for generating the announcement title.

[0063] Step S100: Determine the file to be analyzed from the files that make up the announcement based on the file type.

[0064] Specifically, stock announcements typically consist of multiple documents of different file types. Therefore, based on the file type, a specific document can be selected as the document to be analyzed for subsequent text analysis to obtain a suitable announcement title.

[0065] For example, US stock announcements typically consist of a main file and zero or more supplementary files. The main file is in HTML format, while the supplementary files may be in HTML or other formats. Based on the file type, the main file or a specific supplementary file can be identified as the text to be analyzed, so that subsequent text analysis can be performed on the file to be analyzed to obtain the appropriate announcement title.

[0066] Step S200: Convert the text content of the file to be analyzed into a set of line vectors.

[0067] Specifically, after determining the file to be analyzed, the text content of the file is split into multiple line vectors corresponding to the text content, and a set of line vectors is generated based on these multiple line vectors for subsequent filtering of the line vectors in the set.

[0068] Step S300: Filter the row vectors in the set of row vectors according to font type and layout style to determine the target row vector.

[0069] Specifically, after generating the set of line vectors, it is necessary to filter the line vectors in the set to determine the target line vectors that may be used as the announcement title. Since the line vectors are obtained based on the text content, and some information in the text content is not suitable as the announcement title, and this part of the information is significantly different from other information in the text content in terms of font type and layout style, the line vectors in the set can be filtered according to the font type and layout style to obtain the target line vectors, so as to use the target line vectors as the candidate announcement titles.

[0070] As a feasible implementation example, for instance, in text content, font type and typography style typically follow these rules / principles:

[0071] Font size: Headings are usually larger than the body text.

[0072] Font weight: The font of the title is usually bold.

[0073] Font style: Slanted fonts in text content are often used as notes rather than titles and body text, and therefore are not suitable as announcement titles.

[0074] Text alignment: Text content that uses right alignment is mostly contact information such as phone numbers, addresses, and email addresses, which is not suitable as an announcement title.

[0075] Top margin: If there are multiple headings in the text content, the top margins of each heading are the same, and they are different from the top margins of the rest of the text.

[0076] Tag name: Headings in the text content will not be placed in the tag name 'table'.

[0077] Parent tag name: Text content with a parent tag of 'table' will not be a heading.

[0078] Number of blank lines between text content: The number of blank lines between multi-line headings in the text content is equal, and the number of blank lines between lines of the same node in the text content is small, while the number of blank lines between lines of different nodes is large.

[0079] In order to facilitate subsequent text analysis, word standardization can be performed on the text content, that is, removing word variations and abbreviations and converting them into lowercase forms, thereby unifying the word forms of the text content.

[0080] As one possible implementation, such as Figure 2As shown, the bolded text in the middle of the page is the target line vector to be obtained, i.e., the candidate title. It may be divided into multiple lines, and its font will be larger than the body text, bolded, or enclosed in a 'b' tag. The title may be preceded by redundant information—contact information, some supplementary content, date information, etc. Furthermore, titles are usually not placed in tables, so when a 'table' tag is detected, it can be directly excluded. Phone numbers, email addresses, dates, etc., can be excluded using regular expression matching. Fixed text content containing tags such as 'contact' and 'release', and which is short in length, is also likely not a title and can be excluded. Generally, titles are not right-aligned, so right-aligned text content is also excluded. After the above elimination process, if the text content of the next line is bolded or larger than normal text, it is highly likely to be the candidate title to be obtained.

[0081] Furthermore, since the title in the text content may be divided into multiple lines, the first line cannot be returned as the title alone. It is necessary to determine whether the properties of the following lines are similar to those of the first line. After obtaining the lines with the same properties, they are combined into the same target line vector to serve as the complete title.

[0082] Thus, based on the above rules / principles, the title in the text content can be identified, and the line vector corresponding to the title can be determined as the target line vector, which can then be used as a candidate announcement title.

[0083] Step S400: Generate an announcement title based on the target row vector.

[0084] Specifically, after filtering the row vectors in the row vector set and determining the target row vector as a candidate announcement title, the final announcement title is generated from the candidate announcement titles.

[0085] As a possible implementation, when there is only one candidate announcement title for the target row vector, that candidate announcement title is used as the final announcement title.

[0086] As a feasible implementation, when there are multiple candidate announcement titles for the target row vector, the final announcement title is generated by selecting from the multiple candidate announcement titles through manual review, i.e., in response to the selection instruction triggered by the user.

[0087] As a feasible implementation, in this embodiment, the title of the existing SEC document is used as the default announcement title. If the target row vector cannot be obtained as a candidate announcement title based on the above steps, the default announcement title is used as the final announcement title. If the target row vector can be obtained as a candidate announcement title based on the above steps, the final announcement title is generated based on the candidate announcement title to replace the default announcement title.

[0088] Furthermore, based on the above embodiments, the step S200 of converting the text content of the file to be analyzed into a set of line vectors includes the following steps S201 to S203.

[0089] Step S201: Detect the tag information contained in the file to be analyzed.

[0090] Specifically, different labels are attached to each part of the text content of the file to be analyzed. Therefore, the label information contained in the file to be analyzed can distinguish each part of the text content and determine the algorithm to be used when converting the text content into line vectors.

[0091] Step S202: Based on the tag information, traverse and recursively analyze the text content of the file to be analyzed to convert the text content into multiple line vectors.

[0092] Specifically, the text content of the file to be analyzed is traversed recursively by label to split the text content of the file to be analyzed into multiple parts of text content, and the multiple parts of text content are converted into row vectors accordingly.

[0093] Step S203: Generate a set of row vectors based on the plurality of row vectors.

[0094] Specifically, after converting to obtain multiple row vectors, these multiple row vectors are combined to generate a set of row vectors, which can then be filtered to obtain the target row vector.

[0095] Furthermore, based on the above embodiments, the step S202 above, which involves traversing and recursively analyzing the text content of the file to be analyzed according to the tag information to convert the text content into multiple line vectors, includes the following steps S2021 and S2022.

[0096] Step S2021: Determine the main body of the text content of the file to be analyzed based on the tag information.

[0097] Step S2022: Based on preset segmentation labels and preset table labels, the main body is traversed recursively to convert the text content into multiple row vectors.

[0098] Specifically, text content suitable as an announcement title is typically found in the main body of the document being analyzed. Furthermore, during the conversion of text content into line vectors, the text content needs to be categorized to convert categorized text content belonging to the same category into the same line vector. Additionally, the algorithm used for the line vector conversion process needs to be determined. Segmentation tags and table tags are relatively suitable for categorizing text content.

[0099] As a feasible implementation, for example, the text content of US stock announcement documents typically contains sub-tags 'p' or 'div', as well as the sub-tag 'table'. It can be understood that the sub-tags 'p' or 'div' are used to define paragraph formatting; therefore, text content with only the sub-tags 'p' or 'div' is defined as the same section and converted into a row vector. The sub-tag 'table', on the other hand, is used to define table formatting; therefore, text content with only the sub-tag 'table' requires special processing using table processing functions. Specifically, the table content may contain many nested row tags, and each row tag may contain multiple nested column tags. The special processing involves categorizing the content of all column tags within a single row tag in the table content and converting them into the same row vector.

[0100] In practical application, the first step is to detect the 'body' sub-tag in the text content of the US stock announcement document to determine the corresponding main body. Then, based on the aforementioned sub-tags 'p' or 'div' and 'table', the text content of this main body is recursively traversed. Text content matching the above sub-tags is converted into corresponding line vectors. If it does not match the above sub-tags and there are no other sub-tags to traverse recursively, the remaining text content is directly converted into the same line vector. After converting all the text content of the US stock announcement document into multiple line vectors, these multiple line vectors are combined to generate a corresponding line vector set, which can then be used for subsequent text analysis to select the target line vector.

[0101] Thus, this embodiment provides a specific method for converting the text content of the text to be analyzed into multiple line vectors, and generating a set of line vectors based on the multiple target line vectors, thereby improving the practicality of the technical solution of this application.

[0102] Furthermore, based on the above embodiments, the step S100 above, which determines the file to be analyzed in the files constituting the announcement according to the file type, includes the following steps S101 to S103.

[0103] Step S101: Detect the file type of the files that make up the announcement.

[0104] Step S102: Determine the target attachment file in the files that make up the announcement according to the file type.

[0105] Step S103: The target attachment file is identified as the file to be analyzed.

[0106] As a feasible implementation, for US stock announcements, the documents that make up the US stock announcement include a main document and zero or more supplementary documents. The main document and supplementary documents are different in terms of document type. The main document is usually used to describe the detailed content of a specific event, while the supplementary documents may contain press releases. The title of the press release is more suitable as the title of the announcement. Based on this, the supplementary documents are more suitable as the documents to be analyzed for text analysis to obtain the target line vector. Therefore, based on the document type, the required target supplementary document can be determined from multiple supplementary documents and identified as the document to be analyzed.

[0107] As a possible implementation, for example, an attachment of file type 99.1 is usually a press release for this US stock announcement, so it is highly likely that the attachment contains a suitable announcement title.

[0108] Thus, in this embodiment, the attachment file has a higher priority than the main file in text analysis. At the same time, since the main file and the attachment file are different in file type, the target attachment file can be quickly and accurately determined by detecting the file type, and the target attachment file can be used as the file to be analyzed, which improves the practicality of the technical solution of this application.

[0109] In one feasible implementation, for US stock announcements, if the documents constituting the US stock announcement only include the main file, then the main file is used as the file to be analyzed, and the text content of the main file is converted into lines.

[0110] The vector set is then used to filter the row vectors in the set through text analysis to obtain the target row vector, and finally the announcement title is generated based on the target row vector.

[0111] It is understood that in the above embodiments, when the main file is used as the file to be analyzed, each of the main file's...

[0112] The probability of obtaining a suitable announcement title varies depending on the node text content. Therefore, text analysis can be performed on the text content corresponding to nodes with higher priority based on the node priority list set by the user in advance.

[0113] like Figure 3 As shown, the node item 5.02 in the main file is matched according to the node priority list, and the following lines are used as the target line vector, i.e., the candidate titles; if the content below this node name is a long paragraph, then the first line needs to be matched as the candidate title.

[0114] 5. Further, in a feasible embodiment, the step S102 above, which determines the target attachment file in the file constituting the announcement according to the file type, includes the following steps S1021 to S1023.

[0115] Step S1021: Determine the additional files in the files that make up the announcement based on the file type.

[0116] Step S1022: Compare the attached file with a preset priority list to obtain a comparison result. The preset priority list is used to reflect the probability of obtaining the appropriate announcement title in various types of attached files.

[0117] Step S1023: Determine the target attachment file based on the comparison results.

[0118] As a feasible implementation, the file types of multiple files constituting an announcement are detected, and the attached files are identified. Then, these attached files are compared with a user-preset priority list. That is, different types of attached files have different probabilities of obtaining the appropriate announcement title.

[0119] Users set up priority lists based on practical experience or network data to reflect the probability of obtaining appropriate announcement titles from various types of attachments. By comparing multiple attachments with this priority list, the attachment with the highest / highest priority is selected as the target attachment for subsequent text analysis.

[0120] 0 Further, in a feasible embodiment, the step S102 above, which determines the target attachment file in the file constituting the announcement according to the file type, includes the following steps S1024 to S1026.

[0121] Step S1024: Determine the main file from the files that make up the announcement based on the file type.

[0122] Step S1025: Extract the target attachment information from the text content corresponding to the preset node in the main file.

[0123] Step S1026: Determine the target attachment file based on the target attachment information.

[0124] As a feasible implementation, for US stock announcements, the documents that make up the US stock announcements include a main file and zero or more attachments. A certain node in the main file, such as node Item 9.01, is usually an attachment information table containing information about the attachments. Through this attachment information table, the file type, file name, and other attachment information of each attachment can be extracted, and then the required target attachment file can be determined based on the attachment information for subsequent text analysis of the target attachment file.

[0125] Furthermore, based on the above embodiments, the step S400 of generating the announcement title based on the target row vector includes the following steps S401 to S404.

[0126] Step S401: Convert the text content of the document that makes up the announcement into a preset word set.

[0127] Step S402: Determine the target words in the preset word set that have a preset frequency of occurrence.

[0128] Step S403: Prioritize the target row vectors according to the target words.

[0129] Step S404: Generate the announcement title based on the target row vectors sorted by priority.

[0130] Specifically, the text content of the announcement file, namely the main file and the attached file, is split into a preset word set. Then, the frequency of each word in the preset word set is detected and counted. Words with high frequency can be identified as target words that are more relevant to the content of this announcement. Based on this, some common / meaningless words, such as some conjunctions or auxiliary words, can be removed from the target words. Then, the previously determined target line vectors are prioritized according to the remaining target words. If a target line vector contains a large number of target words, or if the target words in a target line vector have a high frequency of occurrence, the priority of that target line vector is relatively increased. Finally, the announcement title is generated from the target line vectors with higher priority.

[0131] Thus, in this embodiment, by statistically analyzing the frequency of each word in the text content of the document that makes up the announcement, the target word is determined based on the frequency of word occurrence, and then the target line vector is prioritized based on the target word. That is, the target line vector is further filtered based on the frequency of word occurrence, thereby improving the accuracy of the process of selecting the announcement title based on the target line vector.

[0132] The following describes an apparatus embodiment of this application, which can be used to execute the announcement title generation method in the above embodiments of this application. Figure 4A schematic block diagram of the announcement title generation apparatus provided in an embodiment of this application is shown. Figure 4 As shown, the announcement title generation device 400 includes:

[0133] The file determination module 410 is configured to determine the file to be analyzed from the files that make up the announcement based on the file type;

[0134] The text conversion module 420 is configured to convert the text content of the file to be analyzed into a set of line vectors;

[0135] The row vector filtering module 430 is configured to filter the row vectors in the row vector set according to font type and typography style to determine the target row vector;

[0136] The title generation module 440 is configured to generate an announcement title based on the target row vector.

[0137] In one embodiment of this application, based on the above embodiments, the text conversion module includes:

[0138] The detection unit is configured to detect tag information contained in the file to be analyzed;

[0139] The conversion unit is configured to traverse and recursively analyze the text content of the file to be analyzed based on the tag information, so as to convert the text content into multiple line vectors;

[0140] The generation unit is configured to generate a set of row vectors based on the plurality of row vectors.

[0141] In one embodiment of this application, based on the above embodiments, the conversion unit includes:

[0142] The subunit is configured to determine the main body of the text content of the file to be analyzed based on the tag information.

[0143] The transformation subunit is configured to recursively traverse the main body based on preset segment labels and preset table labels to convert the text content into multiple row vectors.

[0144] In one embodiment of this application, based on the above embodiments, the document determination module includes:

[0145] The file determination unit is configured to detect the file type of the files that make up the announcement; determine the target attachment file in the files that make up the announcement according to the file type; and determine the target attachment file as the file to be analyzed.

[0146] In one embodiment of this application, based on the above embodiments, the document determination unit includes:

[0147] The first additional file determination unit is configured to determine additional files from the files constituting the announcement based on file type; compare the additional files with a preset priority list to obtain a comparison result, wherein the preset priority list is used to reflect the probability of obtaining the appropriate announcement title in various types of additional files; and determine the target additional file based on the comparison result.

[0148] In one embodiment of this application, based on the above embodiments, the document determination unit includes:

[0149] The second attachment file determination unit is configured to determine the main file in the files that make up the announcement based on the file type; extract target attachment information from the text content corresponding to the preset node of the main file; and determine the target attachment file based on the target attachment information.

[0150] In one embodiment of this application, based on the above embodiments, the title generation module includes:

[0151] The word conversion unit is configured to convert the text content of the document that makes up the announcement into a preset set of words;

[0152] The word determination unit is configured to determine target words within the preset word set that reach a preset frequency of occurrence.

[0153] A row vector arrangement unit is configured to prioritize the target row vectors according to the target words;

[0154] The title generation unit is configured to generate announcement titles based on priority-sorted target row vectors.

[0155] Figure 5 A schematic block diagram of a computer system architecture for implementing the terminal device of the present application is shown.

[0156] It should be noted that, Figure 5 The computer system 500 of the terminal device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0157] like Figure 5As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 502 or programs loaded from storage section 508 into random access memory (RAM). The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output interface 505 (I / O interface) is also connected to the bus 504.

[0158] The following components are connected to the input / output interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a local area network card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.

[0159] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit 501, it performs various functions defined in the system of this application.

[0160] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0162] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0163] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0164] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0165] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for generating announcement titles, characterized in that, The method for generating the announcement title includes: The file types of the documents that make up the announcement are checked; The target attachment file is determined from the files that make up the announcement based on the file type described; The target attachment file is identified as the file to be analyzed. Convert the text content of the file to be analyzed into a set of line vectors; The row vectors in the set of row vectors are filtered according to font type and typography style to determine the target row vectors; Convert the text content of the document that makes up the announcement into a preset set of words; Identify target words within the preset word set that achieve a preset frequency of occurrence; The target row vectors are prioritized according to the target words; The announcement title is generated based on the target row vectors sorted by priority.

2. The announcement title generation method as described in claim 1, characterized in that, The text content of the file to be analyzed is converted into a set of line vectors, including: Detect the tag information contained in the file to be analyzed; Based on the tag information, the text content of the file to be analyzed is traversed and recursively converted into multiple line vectors; A set of row vectors is generated based on the multiple row vectors.

3. The announcement title generation method as described in claim 2, characterized in that, Based on the tag information, the text content of the file to be analyzed is recursively traversed to convert the text content into multiple line vectors, including: The main body of the text content of the file to be analyzed is determined based on the tag information. The main body is traversed recursively based on preset segmentation labels and preset table labels to convert the text content into multiple row vectors.

4. The announcement title generation method as described in claim 1, characterized in that, Identify the target supplementary files in the files that make up the announcement based on the file type, including: The additional documents are determined from the documents that make up the announcement, based on their document types; The attached file is compared with a preset priority list to obtain the comparison result. The preset priority list is used to reflect the probability of obtaining the appropriate announcement title in various types of attached files. The target attachment file is determined based on the comparison results.

5. The announcement title generation method as described in claim 1, characterized in that, Identify the target supplementary files in the files that make up the announcement based on the file type, including: The main file is determined from the files that make up the announcement based on the file type; Extract the target attachment information from the text content corresponding to the preset node in the main file; The target attachment file is determined based on the target attachment information.

6. An announcement title generation device, characterized in that, The announcement title generation device includes: The file determination module is configured to determine the file to be analyzed from the files that make up the announcement based on the file type; the file determination module includes: a file determination unit, configured to detect the file type of the files that make up the announcement; determine the target attachment file from the files that make up the announcement based on the file type; and determine the target attachment file as the file to be analyzed; The text conversion module is configured to convert the text content of the file to be analyzed into a set of line vectors; The row vector filtering module is configured to filter row vectors in the row vector set according to font type and typography style to determine the target row vector; The title generation module is configured to generate an announcement title based on the target row vector; The title generation module includes: a word conversion unit configured to convert the text content of the document constituting the announcement into a preset word set; a word determination unit configured to determine target words in the preset word set that have a preset frequency of occurrence; a row vector arrangement unit configured to prioritize the target row vectors according to the target words; and a title generation unit configured to generate an announcement title based on the prioritized target row vectors.

7. A terminal device, characterized in that, The terminal device includes: a memory, a processor, and an announcement title generation program stored in the memory and executable on the processor. When the announcement title generation program is executed by the processor, it implements the announcement title generation method as described in any one of claims 1 to 5.

8. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the announcement title generation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Resume data information analyzing processing method, device and equipment, and storage medium

    CN108874928A

  • Announcement text key information extraction method and device

    CN109933796A