Schedule generation method and device, electronic equipment and storage medium
By performing sequence annotation and data structure conversion on event text, the event unit extraction model is trained, which solves the problem that the existing technology is difficult to extract the minimum unit information of nested events or multi-level events, and realizes efficient and accurate event information extraction and schedule generation.
Patent Information
- Application Number
- CN202510196327.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-09
AI Technical Summary
It is difficult for prior art to accurately and completely extract the smallest unit event information in the original event text containing nested events or multi-level events.
By performing sequence annotation of the training text, a first annotation data set of the flat data structure is generated, and converted into a tree data structure, and a second annotation data set containing linear hierarchical structure information is extracted. Then, the event unit extraction model is trained using the second labeled data set, and the minimum unit event information can be extracted automatically and accurately from complex and nested events.
It realizes the rapid, complete and accurate extraction of the smallest unit event information from complex and multi-level nested event message texts, thereby generating a structured event schedule, improving the efficiency and accuracy of event management.
Smart Images

Figure CN119963151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a schedule generation method, device, electronic equipment and storage medium. Background Art
[0002] With the continuous progress and development of society, online events such as online meetings, online speeches, online seminars, display training, etc. have become a common way of communication in many companies and organizations due to their convenience and speed. Efficient and low-cost online events not only break through the limitations of time and space, and do not require venue rental, travel expenses, etc., but also improve the participation and effect of events through interaction, questioning, etc. In addition, the digital transformation of online events is accelerating. With the continuous maturity and application of technologies such as 5G, AI, and big data, online events and hybrid event models will become the mainstream trend.
[0003] Calendar management systems can help companies and organizations better plan and manage online events and improve work efficiency. It has become a trend to display online events in the form of calendars on calendar management systems. For example, Google Calendar has become the first choice for many users due to its seamless integration with the Google ecosystem. Its cross-device synchronization function and smart reminder function greatly facilitate the scheduling of individuals and teams.
[0004] In the era of information explosion, various events are released by different people, at different times and places, and through different channels. It is indeed difficult to quickly and accurately obtain event information from the massive amount of information. Currently, Bert (Bidirectional Encoder Representations from Transformers) is generally used to extract information from events. The Bert model can perform in-depth analysis of texts through its powerful language understanding ability to identify relevant information about events. Specifically, Bert can be used to extract event trigger words and arguments. Trigger words are often verbs and nouns, and the corresponding event arguments are mainly subject, object, time, place, etc.
[0005] However, events often contain nested events or multi-level events. It is difficult to accurately and completely extract the smallest unit of event information based on semantic segmentation alone. This is because semantic segmentation may ignore some details when processing complex text structures, resulting in some parts of the event not being accurately identified. Summary of the invention
[0006] The purpose of the embodiments of the present invention is to provide a schedule generation method, device, electronic device and storage medium, so as to accurately and completely extract the minimum unit event information from the original event text containing nested events and multi-level events.
[0007] To solve the above technical problems, an embodiment of the present invention provides a schedule generation method, comprising: collecting event message text, the event message text including at least one linear hierarchical structure information; inputting the event message text into an event unit extraction model to obtain the minimum unit event information contained in the event message text, the minimum unit event information being event information without sub-level information; generating the schedule according to the minimum unit event information; before inputting the event message text into the event unit extraction model, the method also comprises: performing sequence annotation on the text to be trained to obtain a first annotated data set, the first annotated data set being a flat data structure; converting the first annotated data set into a tree data structure, extracting a second annotated data set containing the linear hierarchical structure information, the second annotated data set being a linear data structure, and the second annotated data set including at least one minimum unit event information of the text to be trained; and performing model training using the second annotated data set to obtain the event unit extraction model.
[0008] An embodiment of the present invention also provides a schedule generation device, comprising: a collection module: used to collect event message text, the event message text includes at least one linear hierarchical structure information; an extraction module: used to input the event message text into an event unit extraction model to obtain the minimum unit event information contained in the event message text, the minimum unit event information is event information without sub-level information; a generation module: used to generate the schedule according to the minimum unit event information; before inputting the event message text into the event unit extraction model, the extraction module is also used to: perform sequence annotation on the text to be trained to obtain a first annotation data set, the first annotation data set is a flat data structure; convert the first annotation data set into a tree data structure, extract a second annotation data set containing the linear hierarchical structure information, the second annotation data set is a linear data structure, and the second annotation data set includes at least one minimum unit event information of the text to be trained; use the second annotation data set to perform model training to obtain the event unit extraction model.
[0009] An embodiment of the present invention also provides an electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the schedule generation method.
[0010] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements the schedule generation method when executed by a processor.
[0011] In an embodiment of the present invention, when training an event unit extraction model, the training text is first sequence-labeled to obtain a first labeled data set with a simple and intuitive flat data structure, and then the first labeled data set is converted into a tree data structure, which can clearly represent the nested relationship and hierarchical structure in the training text, and effectively improve the model's ability to understand event nesting and hierarchical relationships. Next, a second labeled data set is extracted from the tree data structure. The second labeled data set is a linear data structure that is easy to understand and calculate. The data structure conversion process combines the advantages of the clear hierarchy of the tree data structure and the simplicity of the linear structure, providing a flexible and efficient solution for extracting the smallest unit event. Finally, the event unit extraction model is trained based on the second labeled data set with a linear data structure, so that a time unit extraction model that can automatically and accurately extract the smallest unit event information from complex and nested events can be trained. By collecting event message texts including at least one linear hierarchical structure information and inputting it into the event unit extraction model, the minimum unit event information contained in the event message text is obtained. The minimum unit event information can be extracted quickly, completely and accurately from the complex, multi-level nested event message text, and then a schedule is generated based on the minimum unit event information, and an intelligent schedule is established, which facilitates users to organize work more efficiently and improve work efficiency and decision-making quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0013] Figure 1 is a flowchart of a schedule generation method according to an embodiment of the present application; Figure 2 is a schematic diagram of a sequence labeling process according to an embodiment of the present application; Figure 3 is a schematic diagram of a process of extracting a second annotated data set of linear hierarchical structure information according to an embodiment of the present application; Figure 4 is a schematic diagram of a process of establishing a hierarchical mapping table according to an embodiment of the present application; Figure 5 is a schematic diagram of a process for extracting linear hierarchical structure information according to an embodiment of the present application; Figure 6 is a schematic diagram of a process of obtaining the second annotated data set according to the hierarchical structure information according to an embodiment of the present application; Figure 7 is a schematic diagram of a schedule generating device according to an embodiment of the present application; Figure 8 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0014] In the related art, when an event message text is received, schedule information is usually generated for these event message texts to facilitate the management and control of the event. At present, methods such as Bert are usually used to extract event information from the message text, and then segmented through semantic segmentation or other preset rules to obtain relevant information such as the time, place, and participants of the event, and generate a schedule accordingly. However, since methods such as Bert may ignore some detailed information when processing complex text structures, some parts of the event cannot be accurately identified. In addition, although the rule-based method has a high accuracy rate, it relies on a lot of manual labor and is difficult to cover all types of events. Therefore, for some complex events and events with multiple levels of nesting, the existing method of extracting relevant information cannot accurately extract the minimum unit event information in the above-mentioned complex events or multi-level nested events.
[0015] Based on this, an embodiment of the present application provides a schedule generation method, including: collecting event message text, the event message text including at least one linear hierarchical structure information; inputting the event message text into an event unit extraction model to obtain the minimum unit event information contained in the event message text, the minimum unit event information is event information without sub-level information; generating the schedule according to the minimum unit event information; before inputting the event message text into the event unit extraction model, it also includes: performing sequence annotation on the text to be trained to obtain a first annotated data set, the first annotated data set is a flat data structure; converting the first annotated data set into a tree data structure, extracting a second annotated data set containing the linear hierarchical structure information, the second annotated data set is a linear data structure, and the second annotated data set includes at least one minimum unit event information of the text to be trained; using the second annotated data set for model training to obtain the event unit extraction model.
[0016] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, each embodiment of the present application will be described in detail below in conjunction with the accompanying drawings. However, it will be appreciated by those skilled in the art that in each embodiment of the present application, many technical details are proposed in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can also be implemented. The division of the following embodiments is for the convenience of description, and the specific implementation of the present application should not constitute any limitation, and the various embodiments can be combined with each other and referenced on the premise of no contradiction.
[0017] An embodiment of the present application relates to a schedule generation method. The specific flowchart of the schedule generation method is as follows: Figure 1 As shown, the following steps are included: Step S11, collecting event message text, wherein the event message text includes at least one linear hierarchical structure information.
[0018] In an embodiment of the present application, the event may be an online meeting, online training, online speech, online seminar or other activities. In a fast-paced modern work environment, event arrangement is a key link in communication, collaboration, and decision-making. However, relevant information about an event is often scattered in a variety of data sources, such as emails, group messages, conference posters, QR codes, invitations, etc. Effectively integrating these scattered event message texts is essential for generating accurate and comprehensive schedules. In this embodiment, event message texts are collected from a variety of data sources, and a schedule is automatically generated based on the collected event message texts to improve the efficiency and accuracy of processing a large amount of complex event data.
[0019] In one example, when the data source is email, the email content is received through the email protocol, and the email client plug-in automatically identifies event-related keywords in the email, such as "meeting", "activity", "invitation", etc. When an email containing these keywords is detected, the plug-in automatically extracts key information such as the email subject, body content, sender, recipient, and sending time; when the data source is a group message, the API interface of the group message is used to set keyword filtering rules, such as filtering group messages containing keywords such as "project discussion meeting" and "time change". When a group member sends a message in the group chat such as "The project discussion meeting originally scheduled for this Wednesday has been rescheduled to 9 am this Friday. The location remains unchanged, Conference Room 301", the system can promptly capture and extract key information about the change in meeting time; when the data source is a conference poster, the optical character recognition (OCR) technology is used to perform text recognition on the conference poster image, and the text content on the poster is automatically identified, including key information such as the meeting theme, time, and location; when the data source is a QR code, the scanning device sends the link corresponding to the QR code to the server, and the server accesses the link to obtain information such as the time, location, introduction, and guest lineup of the event; when the data source is an invitation letter, the OCR technology is used to convert the image content of the invitation letter into readable text, and the detailed description in the event message is extracted, including the start time, end time, and time schedule of each link, the name of the organizer, the name and position of the invitee, and some precautions.
[0020] In an embodiment of the present application, the scattered event message texts collected from the above-mentioned different data sources are unified and summarized to form an initial message text. The initial message text is often unstructured and contains multiple types of information. In order to efficiently extract and utilize the key information in the event notification text, the initial message text is also classified, especially the text of the time notification type is filtered out. A fine-tuned classification model can be used to realize fully automatic classification and screening of the collected unstructured initial message text to improve the accuracy and efficiency of information processing.
[0021] In an embodiment of the present application, the event message text includes at least one linear hierarchical structure information. In the event message text, especially the notification type text, complex and nested tree data structure information is often included, and these tree data structures imply at least one linear hierarchical structure information, which is crucial for accurately understanding and processing event notifications.
[0022] In the embodiment of the present application, the linear hierarchical structure information is expressed as a hierarchical relationship, which includes multiple levels, and each level has its specific semantics and functions.
[0023] For example, in a meeting notice, the overall arrangement of the meeting may be mentioned first, which includes specific agenda items, and each agenda item may contain sub-agenda or detailed time schedule. These hierarchical structure information are arranged in a certain order, reflecting the process and logic of the meeting. Information between different levels is interrelated, and the information of the previous level usually provides background or framework for the next level. For example, the theme of the meeting is the previous level information, and the specific topic discussion is the next level information under the theme framework. There is a close logical connection between the two.
[0024] For example, suppose there is a meeting notice, and the text of the meeting notice is as follows: [Celebration] Gathering momentum to move forward | Invitation to the 2025 Strategy Meeting of Securities Companies - Chemical Industry Team of Securities Companies —————————————— [Fireworks] New Materials Forum: New Quality Productivity - Investment Opportunities in Chemical Materials #Location: Shanghai Room, 2F, Pujiang Building, Shangri-La Hotel, Shanghai #Time: November 14 (Thursday) 13:30-15:30 #Host: Host 1 13:30-14:10 Frontiers and prospects of robot tactile chemical materials technology 14:10-14:40 Application of coatings industry in automobiles or consumer electronics 14:40-15:10 Application of synthetic biology in the field of pesticides 15:10-15:30 Application of artificial intelligence in the chemical industry —————————————— [Fireworks] Exchange of listed companies in the chemical industry #November 14th afternoon, Shanghai Shangri-La Pujiang Building: #13:30-14:30 Enterprise 1 Meeting Room 710 #14:30-15:30 Table 10, Dalian Changchun Hall, 2nd Floor, Enterprise 2 Enterprise 3 Meeting Room 620 Table 11, Dalian Changchun Hall, 2nd Floor, Enterprise 4 #15:30-16:30 Table 6, Dalian Changchun Hall, 2nd Floor, Enterprise 5 #16:30-17:30 Enterprise 6 Meeting Room 722 —————————————— [Rose] For more details, please contact the chemical team for in-depth discussion! Contact 1 / Contact 2 / Contact 3 / Contact 4 / Contact 5 / Contact 6 / Contact 7 / Contact 8 This meeting notice text contains multiple linear hierarchical structure information. For example, one of the linear hierarchical structure information can be expressed as: First level: overall arrangement of the meeting and first level theme; The second level: the second-level topics nested under the first-level topics, and the overall meeting arrangements corresponding to the second-level topics; The third level: specific matters corresponding to the secondary topics.
[0025] The linear hierarchical structure information in the above meeting notice contains the following content: First level: Gathering momentum to move forward | Invitation letter to the 2025 Strategy Meeting of Securities Company - Chemical Industry Team of Securities Company Contact 1 Second floor: New Materials Forum: New Quality Productivity - Investment Opportunities in Chemical Materials Shanghai Hall, 2F, Pujiang Building, Shangri-La Hotel, Shanghai November 14 (Thursday) 13:30-15:30 Host 1 The third layer: Application of coating industry in automobiles or consumer electronics 14:10-14:40 14:10-14:40 Application of coatings industry in automobiles or consumer electronics Step S12: input the event message text into an event unit extraction model to obtain the minimum unit event information contained in the event message text, wherein the minimum unit event information is event information without sub-level information.
[0026] In the processing of event message text, decomposing complex event information into minimum unit event information is a key step to achieve refined meeting management and efficient schedule scheduling. Minimum unit event information refers to event information without sub-level information, which is the basic element of the entire schedule. In an embodiment of the present application, the minimum unit event information is accurately extracted from the event message text through an event unit extraction model, providing a solid foundation for subsequent schedule generation and management.
[0027] The minimum unit event information is indivisible event information, which represents a specific, independent activity or task in the event, such as a specific topic discussion, a short report session or a break time, etc. These activities cannot be further subdivided into smaller sub-activities.
[0028] Each minimum unit event information usually contains basic event attributes, such as event name, start time, end time, location, participants, etc. These attributes together describe the complete information of the event, allowing it to exist independently and be effectively managed and scheduled. Although the minimum unit event information is independent in structure, there is a logical connection between them, which together constitutes the process of the entire meeting. For example, multiple topic discussion sessions are arranged in a certain order to form the main part of the meeting.
[0029] In the above meeting notification text example, one of the minimum unit event information includes: Theme_1: Gathering momentum to move forward | Invitation letter to the 2025 annual strategy meeting of securities companies - Securities Company Chemical Jin Yiteng Team Topic_1_1: New Materials Forum: New Quality Productivity - Investment Opportunities in Chemical Materials Topic_1_1_1: Application of coatings industry in automobiles or consumer electronics Contact_1: Contact 1 Venue_1_1: Shanghai Hall, 2F, Pujiang Building, Shangri-La Hotel, Shanghai Time_1_1: November 14 (Thursday) 13:30-15:30 Time_1_1_1: 14:10-14:40 Host_1_1: Host 1 Meeting details_1_1_1:14:10-14:40Application of coatings industry in automobiles or consumer electronics In this embodiment, some linear hierarchical structure information corresponding to event attributes has corresponding content starting from the first layer, while some do not start from the first layer, but have corresponding content starting from the second layer or the layers after the second layer. Continuing with the above example, the subject attributes have [Subject_1] [Subject_1_1] [Subject_1_1_1], which have subject content starting from the first layer, while the meeting location has only [Meeting Location_1_1], which has a corresponding meeting location starting from the second layer. For specific event attributes, the starting layer depends on the tree data structure of the event message text. It can be understood that for some event message texts, there is a big subject first, and the big subject does not correspond to a specific event, so there is no place where the event occurs.
[0030] In this embodiment, it is necessary to obtain the acquisition process of the event unit extraction model involved in this step in advance, that is, before the event message text is input into the event unit extraction model, it also includes: The text to be trained is sequence labeled to obtain a first labeled data set, which is a flat data structure; the first labeled data set is converted into a tree data structure, and a second labeled data set containing the linear hierarchical structure information is extracted, the second labeled data set is a linear data structure, and the second labeled data set includes at least one minimum unit event information of the text to be trained; the second labeled data set is used to perform model training to obtain the event unit extraction model.
[0031] By collecting event message text and using the event unit extraction model, we can accurately extract the minimum unit event information from the complex event message text, and then generate a structured event schedule based on this minimum unit event information, thereby improving the efficiency and accuracy of event management. This method can effectively address the shortcomings of traditional methods in dealing with complex events and multi-level nested events.
[0032] Step S13: Generate the schedule according to the minimum unit event information.
[0033] After obtaining the minimum unit event information, the minimum unit event information is displayed on the schedule display interface in the format of the schedule. The schedule display interface can be a schedule function interface in the calendar. Specifically, the minimum unit event information output by the event unit extraction model is stored in a database, and the stored data is presented to the user in an intuitive manner in the calendar user interface (UI). On this interface, the event schedule will be accurately mapped to a specific location on the calendar, so that the user can grasp the future event schedule at a glance.
[0034] In an embodiment of the present application, the schedule generation method can be integrated into applications, such as calendars, mailboxes, memos, etc., and can also be integrated into the internal management system of an enterprise. The event unit extraction model is easy to deploy and can run on mid- and low-end graphics cards, while supporting domestically produced trusted computing graphics cards. After deployment, it can be directly integrated into the systems of various industries through an API interface to automatically obtain upcoming event schedule information, and plan and arrange time in advance accordingly, thereby improving work efficiency and decision-making quality. At the same time, it provides strong support for customers' different query and screening needs, thereby realizing AI empowerment.
[0035] When training the event unit extraction model, the training text is first sequence-labeled to obtain a first labeled data set with a simple and intuitive flat data structure, and then the first labeled data set is converted into a tree data structure, which can clearly represent the nested relationship and hierarchical structure in the training text, and effectively improve the model's ability to understand event nesting and hierarchical relationships.
[0036] Next, the second annotated data set is extracted from the tree data structure. The second annotated data set is a linear data structure that is easy to understand and calculate. The data structure conversion process combines the advantages of the clear hierarchy of the tree data structure and the simplicity of the linear structure, providing a flexible and efficient solution for extracting the smallest unit event.
[0037] Finally, an event unit extraction model is trained based on a second labeled data set with a linear data structure, so that a time unit extraction model that can automatically and accurately extract the minimum unit event information from complex and nested events can be trained.
[0038] By collecting event message texts including at least one linear hierarchical structure information and inputting it into the event unit extraction model, the minimum unit event information contained in the event message text is obtained. The minimum unit event information can be extracted quickly, completely and accurately from the complex, multi-level nested event message text, and then a schedule is generated based on the minimum unit event information, and an intelligent schedule is established, which facilitates users to organize work more efficiently and improve work efficiency and decision-making quality.
[0039] In one embodiment of the present application, the first annotated data set required for training the event unit extraction model includes one or more initial annotations, wherein the initial annotations include entity label information and entity span information; the entity label information includes entity type and level information.
[0040] When processing event message text, in order to train an efficient event unit extraction model, it is necessary to perform sequence annotation on the training text to generate a first annotation dataset containing one or more initial annotations. For example, an initial annotation is: { "id":4668496, "label": "time_1_1_2", "start_offset": 186, "end_offset": 197 } In the above initial annotation, the entity tag information is [time_1_1_2], which includes the entity type, which is "time" here, and the entity tag information also includes the level information, which is "1_1_2" here. It should be noted that in actual applications, another entity tag information marking method can also be used, such as marking the entity type first and then marking the level information, that is, the entity tag information includes two tags, namely [time] and [_1_1_2], where the entity type is [time] and the level information is [_1_1_2].
[0041] The above-mentioned initial annotation also includes entity span information, which may refer to the index range of the starting position and the ending position of a specific event unit in the message text in the event message text, and is used to accurately indicate the position of a specific segment in the text, so as to accurately identify and extract these segments. For example, the text corresponding to the entity tag [time_1_1_2] is 14:10-14:40, which is indicated by the entity span information in the initial annotation, specifically: the starting position start_offset of the entity span information, whose value is 186, indicating the starting character of the text (the position of "1" in 14:10-14:40 in the example), and the ending position end_offset of the entity span information, whose value is 197, indicating the ending character of the above text (the position of the space character after "40" in 14:10-14:40 in the example).
[0042] In another example, the entity span information may also be text content directly obtained through recorded position information, such as 14:10-14:40 in the example.
[0043] In an embodiment of the present application, the annotation system Doccano can be used to perform sequence annotation on the event notification text. Doccano is an open source collaborative annotation tool used to annotate text data for natural language processing (NLP) tasks. It has a friendly interface, can be collaborated by multiple people, and supports a variety of data formats. Figure 2 A detailed introduction to sequence labeling of training texts to generate the first labeled dataset with a flat data structure. Figure 2 The following steps are involved: Step S21, establishing a sequence labeling task.
[0044] Create a sequence annotation task for event information extraction in the Doccano annotation system. To facilitate the management and aggregation of annotation information, you can set a unique task ID for each sequence annotation task. In Doccano, you can distinguish different annotation tasks by project name and description. After the annotation is completed, when exporting the annotation data, each piece of data will contain the task ID to facilitate subsequent data processing and analysis.
[0045] Step S22: define a sequence tag, wherein the sequence tag includes an entity tag and a level tag.
[0046] Each sequence label consists of an entity label and a level label. The entity label corresponds to the entity type, and the level label corresponds to the level information. Taking a meeting as an example, specific entity labels include: meeting details, meeting theme, meeting time, organizer, key points, attending guests, meeting methods, meeting links, meeting phone numbers, meeting numbers, meeting passwords, hosts, speakers, contacts, etc.; hierarchical labels are represented by underscores and Arabic numerals, such as [_1], [_1_1], [_1_1_1], [_1_1_2], [_1_2], [_1_2_1], [_1_2_2], [_2], ..., an underscore plus a number represents the parent level, such as [_1] represents the parent level, and an underscore plus a number after the parent level represents the child level, such as [_1_1] is the child level of [_1], and an underscore plus a number after the child level represents the grandchild level, such as [_1_1_1] is the child level of [_1], and the grandchild level of [_1], and so on.
[0047] Continuing with the above example, if a sequence label is [Topic_1_1], its entity label is marked as Topic, and its level label is marked as [_1_1]. In this way, a sequence label can indicate both the entity type and the level of the marked text in the event. The level of the event is reflected in the sequence label, and level labels with the same prefix indicate that the parent level is the same.
[0048] Step S23: importing the text to be trained into the sequence labeling task.
[0049] In natural language processing (NLP) tasks, importing the text to be trained into the sequence labeling task is a key step in data preparation. This step ensures that the labeling team can begin to label the text in detail and provide high-quality labeled data for subsequent model training.
[0050] Step S24: performing sequence annotation on the text to be trained to obtain a first annotated data set including one or more initial annotations.
[0051] For each line of text, a sequence label is marked according to the semantics and function of the text, such as marking "Preparing for Progress | Invitation Letter to Kaiyuan Securities' 2025 Annual Strategy Conference-Securities Company Chemical Team" as [Topic_1], "New Materials Forum: New Quality Productivity-Chemical Materials Investment Opportunities" as [Topic_1_1], and "#Host: Host 1" as [Host_1_1]. I will not list them one by one here.
[0052] The first annotation data set annotated by the Doccano annotation system is exported to obtain a first annotation data set in a Json array format. The first annotation data set is a flat data structure, which only contains label information and entity span information. The export result is as follows: { "id":4668496, "label": "time_1_1_2", "start_offset": 186, "end_offset": 197 } { "id":4668497, "label": "time_1_1_3", "start_offset": 216, "end_offset": 227 } { "id":4668498, "label": "time_1_1_4", "start_offset": 242, "end_offset": 254 } The sequence tagging result includes one or more initial tags, wherein the initial tags include entity tag information and entity span information; the entity tag information includes entity type and level information.
[0053] In this embodiment, a sequence labeling task for event information extraction is created, in which a training text is sequence labeled to generate a first labeled data set having a flat structure including one or more initial labels, wherein the entity label information in the initial labeling includes entity type and hierarchical information, so as to facilitate subsequent nested hierarchical analysis of the first labeled data set.
[0054] In another embodiment of the present application, the first annotated data set is converted into a tree data structure, such as converting the first annotated data set into a tree data structure through Python code, extracting a second annotated data set containing the linear hierarchical structure information, and the process of extracting the second annotated data set containing the linear hierarchical structure information is as follows: Figure 3 As shown, it includes the following processes: Step S31: establishing a hierarchical mapping table according to the first annotated data set, wherein the hierarchical mapping table includes a mapping relationship between the hierarchical information and the entity span information.
[0055] Constructing a hierarchical mapping table is a key step in understanding the hierarchical relationship between events in the event message text. The hierarchical mapping table records the correspondence between the hierarchical information of the event and the entity span information, which is helpful for subsequent data analysis and model training. The hierarchical mapping table includes the mapping relationship between the hierarchical information and the entity span information. Continuing with the example of the above meeting notice, and taking the entity span information directly as text content as an example, the constructed hierarchical mapping table is shown in the following table:
[0056] Step S32: extracting the linear hierarchical structure information from the hierarchical mapping table, wherein the linear hierarchical structure information includes one or more hierarchical information, and there is a progressive parent-child relationship between the multiple hierarchical information.
[0057] In one example, after constructing the hierarchical mapping table, the next step is to extract linear hierarchical structure information with parent-child relationships from the hierarchical mapping table. This structural information can efficiently understand the hierarchical relationship in the text, such as the relationship between the first-level agenda and the second-level agenda in a meeting notice.
[0058] Continuing with the above meeting notification example, multiple linear hierarchical structure information can be extracted from the hierarchical mapping table, including three linear hierarchical structure information, namely: [_1][_1_1][_1_1_1], [_1][_1_1][_1_1_2], and [_1][_1_2].
[0059] Step S33: Obtain the second labeled data set according to the linear hierarchical structure information.
[0060] The second annotated data set includes at least one minimum unit event information of the text to be trained. Continuing with the above example, taking the linear hierarchical structure information [_1][_1_1][_1_1_1] as an example, a piece of annotated data in the second annotated data set is obtained as follows: Theme_1: Gathering momentum to move forward | Invitation letter to the 2025 annual strategy meeting of securities companies - Securities Company Chemical Jin Yiteng Team Topic_1_1: New Materials Forum: New Quality Productivity - Investment Opportunities in Chemical Materials Topic_1_1_1: Application of coatings industry in automobiles or consumer electronics Contact_1: Contact 1 Venue_1_1: Shanghai Hall, 2F, Pujiang Building, Shangri-La Hotel, Shanghai Time_1_1: November 14 (Thursday) 13:30-15:30 Time_1_1_1:13:30-14:10 Host_1_1: Host 1 Conference details_1_1_1: Frontiers and prospects of robot tactile chemical materials technology In this embodiment, a hierarchical mapping table is used to convert the first annotated data set into a tree data structure with clear parent-child relationships, and a second annotated data set with a linear hierarchical structure is extracted based on the tree data structure, providing structured data support for subsequent model training and application, and helping to improve the model's understanding and processing capabilities of hierarchical relationships in text.
[0061] In another embodiment of the present application, the hierarchical mapping table is established according to the first annotated data set. Figure 4 As shown, it includes the following processes: Step S41, traversing the first annotated data set to obtain the level information of the initial annotation.
[0062] Specifically, the first annotated data set is traversed to obtain entity type and level information in the entity tag information, and the entity type and level information are recorded in a table to obtain the level information therein. The following table is an example of recorded information:
[0063] Step S42, merge one or more entity span information corresponding to the initial annotation with the same hierarchical information to obtain a merged set of entity span information, and obtain the hierarchical mapping table based on the merged set of entity span information, wherein the hierarchical mapping table contains a mapping relationship between the hierarchical information and the merged set of entity span information.
[0064] Specifically, the entity span information with the same hierarchical information is merged and saved in the hierarchical mapping table. Continuing with the example of the above conference notice, taking the hierarchical information [_1_1] as an example, its corresponding entity span information includes "New Materials Forum: New Quality Productivity-Chemical Materials Investment Opportunities", "Shanghai Hall, 2nd Floor, Pujiang Building, Shanghai Shangri-La Hotel", "November 14 (Thursday) 13:30-15:30" and "Host 1". These entity span information are merged together and saved in the hierarchical mapping table, corresponding to the hierarchical information [_1_1]. In the same way, after merging all the hierarchical information and entity span information, the hierarchical mapping table can be obtained.
[0065] In this embodiment, the hierarchical information in the initial annotation is obtained, and based on the hierarchical information, the entity span information corresponding to the same hierarchical information can be quickly and accurately integrated together, and the entity span information corresponding to different hierarchical information can be separated, so as to obtain a hierarchical mapping table corresponding to the hierarchical information and the entity span information, thereby laying the foundation for the subsequent extraction of the hierarchical structure in the event.
[0066] In another embodiment of the present application, the linear hierarchical structure information is extracted from the hierarchical mapping table as follows: Figure 5 As shown, including: The level information in the level mapping table is traversed.
[0067] The hierarchical mapping table includes hierarchical information and entity span information. The hierarchical information in the hierarchical mapping table is traversed to obtain the hierarchical information in the event.
[0068] Determine whether the current level information has any child levels.
[0069] Determine whether the current level is the minimum level according to the level information. If it is the minimum level, obtain the parent of the current level information in turn to obtain the linear level structure information of the current level information. Otherwise, continue to traverse the level information in the level mapping table. For example: Get the level label [_1]. Since the level label [_1] has the child level labels [_1_1] and the level label [_1_2], the current level label [_1] is not the minimum level, so continue to traverse the level mapping table.
[0070] When the obtained hierarchical information is the smallest hierarchical level, such as the current hierarchical level [_1_1_2], which has no child levels, the parents of the current hierarchical information are obtained in sequence, that is, the parent [_1_1] of the current hierarchical information is obtained, and then the parent [_1] of [_1_1] is obtained, and the linear hierarchical structure information of the current hierarchical information is obtained as levels [_1], [_1_1] and [_1_1_2].
[0071] In this embodiment, the linear hierarchical structure information of the event can be extracted conveniently and quickly through the hierarchical information in the hierarchical mapping table.
[0072] In another embodiment of the present application, the process of obtaining the second labeled data set according to the linear hierarchical structure information is as follows: Figure 6 As shown, it includes the following processes: Step S61: determining one or more level information included in the linear level structure information according to the linear level structure information.
[0073] Continuing with the above example, one or more hierarchical information is obtained from the linear hierarchical structure information, for example, the obtained hierarchical information is [_1], [_1_1], and [_1_1_2].
[0074] Step S62: Acquire one or more entity span information corresponding to the level information from the level mapping table.
[0075] Obtain the entity span information corresponding to the hierarchical information from the hierarchical mapping table. For example, the entity span information of the hierarchical information [_1] is obtained as follows: Gathering momentum to move forward | Invitation letter to the 2025 annual strategy meeting of securities companies - Chemical Industry Team of securities companies Contact 1 Step S63: confirm the corresponding entity type according to the entity span information.
[0076] The corresponding entity tag information is confirmed from the initial annotation corresponding to the entity span information. For example, the entity tag information corresponding to the above entity span information is [Topic_1] and [Contact_1]. The entity tag information includes entity type and hierarchical information, and the corresponding entity types [Topic] and [Contact] are obtained.
[0077] Step S64: merge the entity type and the entity span information to obtain the second annotated data set.
[0078] Taking the entity span information as text content directly as an example, according to the determined entity span information and its corresponding entity type, the entity type is first and the entity span information is later merged, separated by a colon. Continuing with the above example, the entity type "topic" and the content corresponding to the entity span information are merged to obtain "Topic: Poised to move forward | Securities Company 2025 Annual Strategy Meeting Invitation Letter - Securities Company Chemical Team" and "Contact: Contact 1". In the same way, after merging all the content, the second labeled data set can be obtained.
[0079] In this embodiment, one or more hierarchical information protected by the hierarchical structure information is determined according to the hierarchical structure information, so that all hierarchical levels in the hierarchical nested event can be accurately obtained to avoid hierarchical omissions and confusion. Then, one or more entity span information corresponding to the hierarchical information and the entity type corresponding to the entity span information are obtained from the hierarchical mapping table, and the entity type and the entity span information are merged to extract the second annotation data set contained in all hierarchical levels in the event information.
[0080] In another embodiment of the present application, the process of merging the entity type and the entity span information includes the following process: determining whether there are multiple entity span information with the same entity type; if there are multiple entity span information with the same entity type, obtaining the hierarchical information corresponding to the entity span information, and merging them in sequence according to the hierarchical information in the order of parent taking precedence over child.
[0081] Get the entity label information corresponding to the entity span information, and extract the entity type from the entity label information. For example, extract the entity type as "Theme" from the entity span information "Preparing for Momentum | Invitation to Securities Company's 2025 Strategy Meeting-Securities Company Chemical Team", and extract the entity type as "Theme" from the entity span information "New Materials Forum: New Quality Productivity-Chemical Materials Investment Opportunities". The entity types corresponding to the two entity span information are the same. If the entity types are the same, perform the following steps.
[0082] When the judgment result is that the entity types corresponding to the entity span information are the same, the hierarchical information corresponding to the entity span information is obtained. For example, the entity type extracted from the entity span information "Preparing for Progress | Invitation Letter for the 2025 Strategy Meeting of Securities Companies - Chemical Team of Securities Companies" is [Topic], and the entity type extracted from the entity span information "New Materials Forum: New Quality Productivity - Investment Opportunities in Chemical Materials" is [Topic], and the entity type extracted from the entity span information "Application of the Coating Industry in Automobiles or Consumer Electronics" is [Topic]. The entity types corresponding to the three entity span information are the same, and the hierarchical information corresponding to the entity span information "Preparing for Progress | Invitation Letter for the 2025 Strategy Meeting of Securities Companies - Chemical Team of Securities Companies" is obtained [_1], and the hierarchical information corresponding to the entity span information "New Materials Forum: New Quality Productivity - Investment Opportunities in Chemical Materials" is obtained [_1_1], Get the hierarchical information [_1_1_2] corresponding to the entity span information "Application of coating industry in automobiles or consumer electronics", and merge them in order of parent priority over child according to the hierarchical information. Since [_1] is the parent of [_1_1], and [_1_1] is the parent of [_1_1_2], the parent is merged first, and the entity span information corresponding to the hierarchical information [_1] is first, followed by the entity span information corresponding to the hierarchical information [_1_1], and finally the entity span information corresponding to the hierarchical information [_1_1_2]. Use the same merging method to merge all the data to obtain the above-mentioned second labeled data set.
[0083] In an embodiment of the present application, entity span information of the same entity type is merged together in sequence in the order of parent taking precedence over child, so that the entity span information structure corresponding to each entity type is clear, which facilitates the subsequent training of the event unit extraction model based on the second labeled data set.
[0084] When training the event unit extraction model, the basic large model can choose llama3.1-8B-chinese-chat, qwen2.5-7B-instruct, and the fine-tuning method can use the low-rank adaptation method (Lora) and the efficient full parameter fine-tuning method (BAdam). In the embodiment of the present application, qwen2.5-7B-instruct is used for Lora fine-tuning, and the improved low-rank adaptation method Lora+ and rsLora are used to improve the fine-tuning effect, and the fine-tuned model is quantized by gptq (Gradient-based Post-training Quantization) to reduce the hardware requirements while meeting the accuracy.
[0085] After the event unit extraction model is trained, the event notification text is recognized using the event unit extraction model to obtain the event information of the smallest unit, which is then stored in a database for further processing and display.
[0086] In the embodiment of the present invention, the minimum unit event information is extracted from the event message text through the event unit extraction model, so that the minimum unit event information can be extracted quickly, completely and accurately from the complex and multi-level nested event message text.
[0087] The steps of the above method are divided only for clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of the present invention. Adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the protection scope of the invention.
[0088] In addition, the examples mentioned in the above embodiments can be freely combined, and any combination can be understood as an embodiment. The "embodiment" or "example" appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It can be understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0089] Another embodiment of the present invention relates to a schedule generating device, such as Figure 7 As shown, the schedule generating device 100 includes: Collection module 101: used for collecting event message text, wherein the event message text includes at least one linear hierarchical structure information.
[0090] Extraction module 102: used for inputting the event message text into an event unit extraction model to obtain the minimum unit event information contained in the event message text, wherein the minimum unit event information is event information without sub-level information.
[0091] Generating module 103: used to generate the schedule according to the minimum unit event information.
[0092] Before inputting the event message text into the event unit extraction model, the extraction module 102 is also used to: perform sequence annotation on the text to be trained to obtain a first annotated data set, which is a flat data structure; convert the first annotated data set into a tree data structure, and extract a second annotated data set containing the linear hierarchical structure information, which is a linear data structure, and includes at least one minimum unit event information of the text to be trained; and use the second annotated data set to perform model training to obtain the event unit extraction model.
[0093] In some embodiments, the first annotation data set includes one or more initial annotations, wherein the initial annotations include entity tag information and entity span information; the entity tag information includes entity type and level information.
[0094] In some embodiments, the step of performing sequence labeling on the text to be trained to obtain a first labeling data set includes: establishing a sequence labeling task; defining sequence labels, wherein the sequence labels include entity labels and hierarchical labels; importing the text to be trained into the sequence labeling task; performing sequence labeling on the text to be trained to obtain a first labeling data set including one or more initial labelings; the entity type corresponds to the entity label, and the hierarchical information corresponds to the hierarchical label.
[0095] In some embodiments, converting the first annotated data set into a tree data structure and extracting a second annotated data set containing the linear hierarchical structure information includes: establishing a hierarchical mapping table based on the first annotated data set, the hierarchical mapping table including a mapping relationship between the hierarchical information and the entity span information; extracting the linear hierarchical structure information from the hierarchical mapping table, the linear hierarchical structure information including one or more of the hierarchical information, and there is a progressive parent-child relationship between the multiple hierarchical information; and obtaining the second annotated data set based on the linear hierarchical structure information.
[0096] In some embodiments, establishing a hierarchical mapping table based on the first annotated data set includes: traversing the first annotated data set to obtain the hierarchical information of the initial annotation; merging one or more entity span information corresponding to the initial annotation with the same hierarchical information to obtain a merged set of entity span information, and based on the merged set of entity span information, obtaining the hierarchical mapping table, wherein the hierarchical mapping table contains a mapping relationship between the hierarchical information and the merged set of entity span information.
[0097] In some embodiments, extracting the linear hierarchical structure information from the hierarchical mapping table includes: traversing the hierarchical information in the hierarchical mapping table; determining whether the current hierarchical information has any children; and if the hierarchical information does not have any children, obtaining the parents of the current hierarchical information in turn to obtain the linear hierarchical structure information described by the current hierarchical information.
[0098] In some embodiments, obtaining the second annotated data set based on the linear hierarchical structure information includes: determining one or more of the hierarchical information contained in the linear hierarchical structure information based on the linear hierarchical structure information; obtaining one or more of the entity span information corresponding to the hierarchical information from the hierarchical mapping table; confirming the corresponding entity type based on the entity span information; and merging the entity type and the entity span information to obtain the second annotated data set.
[0099] In some embodiments, the merging of the entity type and the entity span information includes: determining whether there are multiple entity span information with the same entity type; if there are multiple entity span information with the same entity type, obtaining the hierarchical information corresponding to the entity span information, and merging them in sequence according to the hierarchical information in the order of parent taking precedence over child.
[0100] It is not difficult to find that this embodiment is a device embodiment corresponding to the above method embodiment, and this embodiment can be implemented in conjunction with the above method embodiment. The relevant technical details mentioned in the above method embodiment are still valid in this embodiment, and in order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the above method embodiment.
[0101] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of the present invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other units in this embodiment.
[0102] Another embodiment of the present application provides an electronic device, see Figure 8 , including: a processor 201, a memory 202, and a computer program stored in the memory 202 and executable on the processor 201, such as a data processing program. When the processor 201 executes the computer program, the steps in the above-mentioned schedule generation method embodiments are implemented, such as Figure 1 Steps S11-S13 are shown.
[0103] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 202 and executed by the processor 201 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the electronic device.
[0104] The electronic device may be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The electronic device may include, but is not limited to, a processor 201 and a memory 202. Those skilled in the art will appreciate that the schematic diagram is merely an example of an electronic device and does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the diagram, or may combine certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0105] The processor 201 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor 201 may also be any conventional processor, etc. The processor 201 is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire electronic device.
[0106] The memory 202 can be used to store the computer program and / or module. The processor 201 realizes various functions of the electronic device by running or executing the computer program and / or module stored in the memory 202 and calling the data stored in the memory 202. The memory 202 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 202 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0107] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor 201. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0108] Another embodiment of the present application provides a computer-readable storage medium, which stores a program or instruction. When the program or instruction is executed by a processor, the various processes of the embodiment of the schedule generation method as described above are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0109] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0110] Another embodiment of the present application further provides a computer program product, including computer instructions, which, when executed by a processor, implement the various processes of any of the above-mentioned schedule generation method embodiments and can achieve the same technical effect. To avoid repetition, they are not described here.
[0111] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0112] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, storage and other aspects of user personal information involved in the technical solution of the present invention are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information and network security.
[0113] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0114] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for enabling a service classification device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0115] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A schedule generation method, characterized in that: include: Collecting an event message text, wherein the event message text includes at least one linear hierarchical structure information; Inputting the event message text into an event unit extraction model to obtain minimum unit event information contained in the event message text, wherein the minimum unit event information is event information without sub-level information; Generate the schedule according to the minimum unit event information; Before inputting the event message text into the event unit extraction model, the method further includes: Perform sequence labeling on the training text to obtain a first labeling data set, wherein the first labeling data set is a flat data structure; Converting the first annotated data set into a tree data structure, extracting a second annotated data set containing the linear hierarchical structure information, wherein the second annotated data set is a linear data structure, and the second annotated data set includes at least one minimum unit event information of the text to be trained; The event unit extraction model is obtained by performing model training using the second annotated data set.
2. The schedule generation method according to claim 1, characterized in that: The first annotation data set includes one or more initial annotations, wherein the initial annotations include entity tag information and entity span information; the entity tag information includes entity type and level information.
3. The schedule generation method according to claim 2, characterized in that: The step of performing sequence labeling on the training text to obtain a first labeling data set includes: Establish a sequence labeling task; Defining a sequence tag, wherein the sequence tag includes an entity tag and a level tag; Importing the text to be trained into the sequence labeling task; Performing sequence annotation on the text to be trained to obtain a first annotated data set including one or more initial annotations; The entity type corresponds to the entity tag, and the level information corresponds to the level tag.
4. The schedule generation method according to claim 2, characterized in that: The converting the first annotated data set into a tree data structure and extracting a second annotated data set containing the linear hierarchical structure information comprises: Establishing a hierarchical mapping table according to the first annotated data set, wherein the hierarchical mapping table includes a mapping relationship between the hierarchical information and the entity span information; Extracting the linear hierarchical structure information from the hierarchical mapping table, wherein the linear hierarchical structure information includes one or more hierarchical information, and there is a progressive parent-child relationship between the multiple hierarchical information; The second labeled data set is obtained according to the linear hierarchical structure information.
5. The schedule generation method according to claim 4, characterized in that: The establishing of a hierarchical mapping table according to the first annotated data set comprises: Traversing the first annotated data set to obtain the level information of the initial annotation; One or more entity span information corresponding to the initial annotations with the same hierarchical information are merged to obtain a merged set of entity span information. Based on the merged set of entity span information, the hierarchical mapping table is obtained, and the hierarchical mapping table contains a mapping relationship between the hierarchical information and the merged set of entity span information.
6. The schedule generation method according to claim 4, characterized in that: The extracting the linear hierarchical structure information from the hierarchical mapping table comprises: Traversing the level information in the level mapping table; Determine whether the current level information has a sub-level; In the case that the hierarchical information does not have a child level, the parent levels of the current hierarchical information are acquired in sequence to obtain the linear hierarchical structure information of the current hierarchical information.
7. The schedule generation method according to claim 4, characterized in that: The obtaining the second annotated data set according to the linear hierarchical structure information comprises: Determine, according to the linear hierarchical structure information, one or more hierarchical information included in the linear hierarchical structure information; Acquire one or more entity span information corresponding to the level information from the level mapping table; Confirming the corresponding entity type according to the entity span information; The entity type and the entity span information are combined to obtain the second annotated data set.
8. The schedule generation method according to claim 7, characterized in that: The merging of the entity type and the entity span information comprises: Determine whether there are multiple entity span information with the same entity type; In the case that there are a plurality of entity span information with the same entity type, the hierarchical information corresponding to the entity span information is obtained, and the entity span information is merged in sequence in the order of parent taking precedence over child according to the hierarchical information.
9. A schedule generating device, characterized in that: include: Collection module: used for collecting event message text, wherein the event message text includes at least one linear hierarchical structure information; Extraction module: used for inputting the event message text into the event unit extraction model to obtain the minimum unit event information contained in the event message text, wherein the minimum unit event information is event information without sub-level information; Generating module: used for generating the schedule according to the minimum unit event information; Before inputting the event message text into the event unit extraction model, the extraction module is further used to: perform sequence annotation on the text to be trained to obtain a first annotated data set, wherein the first annotated data set is a flat data structure; convert the first annotated data set into a tree data structure, and extract a second annotated data set containing the linear hierarchical structure information, wherein the second annotated data set is a linear data structure, and the second annotated data set includes at least one minimum unit event information of the text to be trained; The event unit extraction model is obtained by performing model training using the second annotated data set.
10. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the schedule generation method according to any one of claims 1 to 8.
11. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the schedule generation method according to any one of claims 1 to 8 is implemented.