SOP generation method and system based on large language model and storage medium
By using real-time audio stream processing based on a large language model and human-machine collaborative review, the inefficiency and lag of traditional SOP generation methods have been solved, enabling instant and accurate SOP generation and continuous optimization, and reducing operating costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional SOP generation methods are inefficient, result in the loss of tacit knowledge, are slow to update, and have non-standard content, leading to long knowledge production cycles, high costs, and difficulty in guaranteeing quality.
By employing real-time audio stream conversion, intent recognition, information extraction, and structured processing based on a large language model, combined with human-machine collaborative review, SOPs can be generated instantly and continuously optimized.
It enables the instant generation and updating of knowledge, ensuring the accuracy and consistency of SOPs, reducing operating costs, and improving the reusability of knowledge and the company's responsiveness.
Smart Images

Figure CN121787402A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of knowledge management and artificial intelligence technology, and in particular to a method, system and storage medium for generating SOPs based on a large language model. Background Technology
[0002] In enterprise operations management, Standard Operating Procedures (SOPs) are the cornerstone for ensuring service quality, improving work efficiency, and achieving scalable expansion. However, despite their crucial importance, the traditional methods of creating and maintaining SOPs have many problems: 1. Inefficient knowledge acquisition: Traditional SOP development relies heavily on manual interviews, meetings, and document preparation. This process is time-consuming and labor-intensive, requiring deep involvement of senior business experts, resulting in long knowledge production cycles and high costs.
[0003] 2. Significant Loss of Tacit Knowledge: A wealth of "living knowledge" and best practices for solving real-world problems arise from real-time interactions between frontline employees and customers or colleagues. This verbal knowledge is often forgotten after the call ends, lacking an effective real-time capture mechanism, resulting in valuable experience being unable to be accumulated and reused. Trying to dig it out by listening to massive amounts of recordings afterward is like looking for a needle in a haystack.
[0004] 3. SOP updates are severely delayed: When a new business problem or solution first appears in a call, the traditional SOP update process requires a lengthy process of manual discovery, summarization, writing, and review, which causes the SOPs in the knowledge base to always lag behind the actual situation on the front line and make it impossible to respond in time.
[0005] 4. Subjective and non-standard content: When SOPs are manually written, the quality and level of detail are limited by the writer's personal experience and expression habits, which can easily lead to omissions of key steps, unclear descriptions, or inconsistent formats, making it difficult to guarantee the standardization and high quality of SOPs. Summary of the Invention
[0006] In view of this, the purpose of this invention is to propose a method for generating SOPs based on a large language model, comprising the following steps: S1. Acquire the audio or text stream of the session in real time, convert the audio stream into a text stream, and preprocess the text stream; S2. Perform real-time intent recognition on the preprocessed text stream using a large language model; S3. When an intention involving problem solving or operation instructions is identified, the context of the preprocessed text stream is analyzed and information is extracted. S4. Using incremental database updates, the analyzed and refined information is mapped in real time to preset structured fields to form structured statements. S5. Organize and integrate the structured fields to obtain the standard operating procedure.
[0007] Preferably, converting a call audio stream into a text stream includes: In multi-party calls, the system identifies the different speakers in real time and uses an automatic speech recognition model that has been fine-tuned on datasets that are specific to business domains and include common dialect accents for speech recognition and conversion.
[0008] Preferably, in step S1, the preprocessing includes: Segment the continuous text stream into meaningful semantic segments or process it according to a fixed time window; Denoising and anonymizing of sensitive information are performed on the text stream; The text stream is normalized, spoken language is converted into written language, and terms and abbreviations are expanded and annotated.
[0009] Preferably, analyzing and extracting information from the context of the current session includes: The core entities in the conversation are marked, and the subjects and objects of operation omitted in the dialogue are inferred based on these core entities, so as to extract key information.
[0010] Preferably, the structured fields are sorted and integrated to obtain the standard operating procedure, which includes: By integrating fragmented structured fields through a large language model, and obtaining the sequence of operation steps based on the timeline and logical dependencies, and supplementing the context, a standard operating procedure is obtained.
[0011] Preferably, this method further includes: After obtaining the standard operating procedures, they are manually reviewed and then released.
[0012] Preferably, this method further includes: The manually modified records generated from the manual review are collected as high-quality supervised learning samples. The large language model was fine-tuned using supervised learning samples.
[0013] This invention also provides a SOP generation system based on a large language model, including a data input and preprocessing module and an SOP generation engine: The data input and preprocessing module is used to capture the audio stream of an ongoing call or the text stream in real time, convert the audio stream into a text stream, preprocess the text stream, and perform intent recognition on the preprocessed text stream using a large language model to determine whether there is a conversation involving problem solving or operation instruction intent, thus triggering the SOP generation engine. When the SOP generation engine is triggered, the large language model processing engine is started to analyze and extract information from the context of the current session. The extracted information is mapped to the preset structured fields in real time through incremental database updates to form structured statements. The structured fields are then sorted and integrated to obtain the standard operating procedure.
[0014] Preferably, this system also includes a human-machine collaborative review and release module and a closed-loop optimization and model iteration module; The human-machine collaborative review and release module is used to obtain the standard operating procedures, submit the output standard operating procedures to human review, and release them after review. The closed-loop optimization and model iteration module is used to collect manually modified records as high-quality supervised learning samples, and to fine-tune the large language model using supervised learning samples.
[0015] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by one or more processors, causes the one or more processors to perform the steps in the above-described SOP generation method based on a large language model.
[0016] The beneficial effects of this invention are: 1. This invention revolutionizes the traditional SOP production cycle. Traditionally, developing an SOP from problem identification to final release requires interviews, data compilation, drafting, and review, taking days or even weeks. This invention compresses this cycle to minutes after a call or conversation ends, enabling "instantaneous knowledge generation and use." Furthermore, this invention automatically generates a draft SOP and sends it for manual review during the "golden moment" when the person involved (such as a customer service expert) has the clearest memory and most complete details, greatly ensuring the completeness and accuracy of the knowledge and avoiding information omissions and biases caused by retrospective recall.
[0017] 2. This invention can accurately capture the "tacit knowledge" and "best practices" that usually exist in the minds of senior employees and are improvised when solving sudden or complex problems in real time, and solidify them into replicable and disseminable explicit knowledge assets the moment they are generated.
[0018] 3. For new product features, new market policies, or new customer issues, this invention can generate a first version of the SOP (Standard Operating Procedure) in the first successfully resolved call. This ensures that the enterprise knowledge base is no longer lagging behind the front-line business, but is fully synchronized with its dynamics.
[0019] 4. This invention continuously learns from the latest sessions, dynamically identifying outdated aspects of existing SOPs or better solutions, enabling high-frequency and rapid iteration of the knowledge base. This ensures that frontline employees always use the latest and most effective work instructions.
[0020] 5. This invention is based on a unified large language model and standard template generation, ensuring consistency in format, style, and logical structure across all SOPs. Simultaneously, through human-machine collaborative review and feedback-based closed-loop optimization, the professionalism and accuracy of the SOPs are doubly guaranteed, and the model's capabilities will continue to improve, resulting in increasingly higher quality drafts generated by AI.
[0021] 6. This invention can significantly reduce enterprise operation and management costs, greatly reduce the time required for business experts, knowledge managers, and trainers, and significantly reduce the human resource costs of knowledge production. The rapidly generated SOPs can immediately serve as training materials for new employees or as an empowerment tool for all employees, shortening the popularization cycle of new skills and processes, reducing repetitive errors and customer complaints caused by knowledge gaps or insufficient skills, thereby reducing overall operating costs. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the method for quickly generating SOPs based on mining conversation content using a large language model, according to an embodiment of the present invention. Figure 2 This is a system block diagram for rapidly generating SOPs based on mining session content using a large language model, as described in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0025] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0026] Example 1: like Figure 1 As shown in the embodiments of this specification, a method for generating SOPs based on a large language model is provided, which specifically includes the following steps: S1. Acquire the audio or text stream of the session in real time, convert the audio stream into a text stream, and preprocess the text stream: This method adapts to interface protocols from different sources, enabling real-time acquisition of audio streams from ongoing calls or real-time text streams from online customer service or communication tools (such as IM tools). The acquired audio or text streams are dynamically generated and continuously transmitted data accompanying the communication activity, rather than static data files. After acquiring the data streams, this method performs low-latency processing on the real-time incoming data, primarily including: If the acquired data stream is a call audio stream, then real-time speech-to-text (Real-time ASR) technology is used to quickly and continuously convert the real-time audio stream into text and achieve streaming output. For multi-party call scenarios, real-time speaker separation is performed, that is, in multi-party calls, the speaker's role corresponding to each segment of speech is labeled in real time, clearly distinguishing the source of speech in multi-party calls, and providing clear role dimension support for subsequent data analysis. Furthermore, an ASR model that has been fine-tuned on datasets with specific business domains (such as IT and customer service) and containing common dialect accents is used to improve the model's speech recognition accuracy and semantic understanding adaptability in target scenarios.
[0027] As one implementation method, the fine-tuning of the common dialect accent dataset in the above-mentioned ASR model includes: targeted collection or annotation of speech data containing common dialect accents (such as dialect accents mixed in with Mandarin) in the main regions covered by the business. The collection scope focuses on the accent types that frequently occur in actual communication, especially speech content in Mandarin mixed with various dialect characteristics, to ensure that the dataset can accurately match the language usage scenarios of the target region. The model training starts with the ASR basic model pre-trained on massive amounts of general Mandarin data, which already has mature general speech recognition capabilities.
[0028] Next, we entered the fine-tuning phase, using a previously constructed exclusive dialect accent dataset to specifically optimize the base model. Key optimizations included: iterative upgrades to the acoustic model within the base model, training it with extensive dialect accent data to allow the model to deeply learn and adapt to the unique pronunciation patterns of dialects, including tonal variations, vowel variations, and initial consonant differences; and adaptation adjustments to the language model within the base model, incorporating specific vocabulary, common expressions, and unique grammatical habits that may be mixed in with dialects to optimize the model's understanding and recognition capabilities of these linguistic phenomena, ultimately achieving a significant improvement in the model's recognition accuracy in dialect accent scenarios.
[0029] As one implementation method, to further improve the accuracy of speech recognition in specific scenarios, this method also includes acoustic model optimization based on hot words. Hot words can be pre-imported specific terms that frequently occur in business scenarios, or relevant vocabulary can be dynamically supplemented according to actual usage needs. The coverage includes industry-specific terms, core product names, and specific expressions commonly used in communication. For the added hot words, targeted optimization of the acoustic model is carried out, taking into account the pronunciation characteristics in dialect environments, and focusing on adjusting the model's recognition logic for hot word pronunciation. By strengthening the feature matching of hot words in dialect pronunciation patterns, the recognition accuracy of hot words themselves is improved, and vocabulary confusion caused by dialect accents can be further avoided.
[0030] As one implementation, the above preprocessing also includes performing text segmentation and windowing on the acquired real-time text stream or the text stream obtained from speech conversion, as well as real-time text cleaning and desensitization.
[0031] Specifically, text segmentation and windowing processing includes: on the one hand, segmenting texts into meaningful complete segments based on semantic logic to ensure that text units conform to natural language expression habits; on the other hand, it also supports regular division according to fixed time windows. Both methods can flexibly adapt to the real-time analysis rhythm of subsequent models, keeping the processing process synchronized with the data generation speed and improving overall analysis efficiency.
[0032] Real-time text cleaning and desensitization includes: 1) Dynamically perform text noise filtering and sensitive information anonymization, automatically remove redundant interjections, meaningless filler content and other irrelevant noise from the text, and at the same time perform compliant anonymization of sensitive information such as mobile phone numbers, names and addresses to improve security.
[0033] 2) Spoken to Written Language Conversion: After completing the ASR text output and before sending it to the Large Language Model (LLM) for SOP extraction, text normalization is initiated to convert colloquial words and phrases such as "can't handle" and "no hope" into formal written expressions such as "cannot be solved" and "failed." This normalization process can be implemented using a lightweight LLM or a rule engine to make the text format more suitable for the processing characteristics of LLM and improve the accuracy of subsequent SOP extraction.
[0034] 3) Automatically identify industry-specific abbreviations in text, such as "ITSM system" and "KPI", and provide full meaning expansion or clear annotation in the context to eliminate ambiguity caused by abbreviations, help LLM accurately grasp the core information of the text, and ensure the effectiveness of subsequent processing.
[0035] S2. Real-time intent recognition is performed on the preprocessed text stream using a large language model, specifically including: Real-time Intent Recognition and State Tracking: This method does not process all conversations. Instead, it utilizes a large language model to perform real-time intent recognition on the conversation stream. It continuously tracks the conversation state, and when it detects a typical problem-solving pattern that matches "problem introduction, analysis and exploration, solution proposal, and verification and resolution," or when a clear "operation instruction" intent appears, the system automatically marks that conversation segment as a potential SOP generation event. This targeted filtering mechanism avoids redundant processing of invalid data and accurately identifies core conversation segments with SOP extraction value, laying the foundation for the subsequent efficient generation of standardized processes.
[0036] S3. When an intention involving problem-solving or operational instructions is identified, the context of the preprocessed text stream is analyzed and information is extracted, specifically including: SOP capture session initiation: Once triggered, the system prioritizes computing resources and analytical power for the current session processing stage. It not only focuses on real-time dialogue content but also analyzes historical interaction information to accurately extract core information related to SOP generation, such as the core issues, key processing steps, decision-making basis, and execution points. After the entire session concludes, a human will intervene to comprehensively verify the information extracted by the system, checking the accuracy of the content, the completeness of the process, and the coherence of the logic. Once confirmed correct, the final submission is completed, providing high-quality foundational material for the subsequent standardization and consolidation of SOPs.
[0037] S4. Using incremental database updates, the analyzed and refined information is mapped in real time to preset structured fields to form structured statements, specifically including: Dynamic element filling drives the large language model by constructing a specific set of system prompts. The core design of these prompts lies in establishing a complete workflow for stateful context processing and iterative knowledge construction for the model. It forces the model to re-examine its understanding of the core elements of the Standard Operating Procedure (SOP) based on the complete historical dialogue after each round of conversation, and to incrementally update (fill, add, correct) the refined information into pre-defined structured fields in real time. For example, when a user describes a problem in fragmented ways during multiple rounds of conversation, the system continuously summarizes this scattered information, gradually refining and filling it into the "Problem / Scenario Description" field, ensuring that this field always maintains a complete and accurate summary of the problem. Meanwhile, the operational guidance provided by customer service during the conversation is broken down into atomic steps and added one by one to the "Operation Steps" list, enabling unstructured dialogue content to be dynamically transformed into well-organized, structured SOP knowledge.
[0038] In its implementation, this step requires maintaining a "SessionEntityState" as an anchor point for contextual understanding. When a user mentions "My server is down," the system marks "Server A" as the current core entity. In subsequent conversations, when the expert says "Please restart it," the system infers through "stateful context processing" that the omitted subject is "you (the customer)," and the object of "restarting" is "it (Server A)," thus avoiding information gaps caused by omissions or ambiguous references in spoken language. This ensures that the filling of structured fields remains consistent with the dialogue logic, achieving a precise transformation from fragmented interaction to structured knowledge.
[0039] S5. Organize and integrate the resulting structured fields to obtain the standard operating procedure, which specifically includes: Context integration and logical organization: At the critical juncture of the call ending or problem resolution, this step integrates all fragmented information extracted in real time from the large model during the conversation. This fragmented information includes scattered operation prompts, temporary analysis conclusions, conditional constraints, user feedback, etc., which need to be organized into an organic whole through multi-dimensional analysis.
[0040] Specifically, the system first establishes a complete timeline based on the chronological order of the conversations, precisely connecting key milestones such as the initial mention of the problem, preliminary investigation, solution discussion, execution, and result verification, ensuring a clear and traceable timeline for each step. Second, it deeply identifies and organizes the logical dependencies in the resolution process, clarifying that a certain step depends on the completion of another, or that a particular analytical conclusion is based on specific feedback, thus avoiding logical gaps caused by fragmented information. Simultaneously, for instances of omissions or ambiguous references due to colloquial speech habits, the system supplements the context with complete background information, ensuring that the preconditions and implementation scenarios for each step are clear and explicit. Based on this, the system automatically streamlines the sequence of operational steps according to logical dependencies and the timeline, correcting any possible reversals or repetitions to ensure the chain of steps is practically executable.
[0041] Through this series of integrations and sorting, the originally scattered information will be woven into a logically coherent and complete knowledge system, providing a structured and highly usable basic framework for the subsequent standardization, refinement, and consolidation of SOPs.
[0042] As one implementation method, this method further includes: S6. Human-Machine Collaborative Auditing and Model Iteration: It provides a visual review interface that intuitively presents AI-generated SOP knowledge points in real time, including structured problem descriptions, operation steps, and core elements. This allows frontline agents or domain experts to immediately proofread and modify the content after a call. At this point, the agent or expert has a clear memory of the call details, problem background, and solutions, enabling them to quickly identify potential issues such as expression bias, missing steps, logical inconsistencies, or inaccurate information in the AI-generated content. During the review process, content correction, key information supplementation, and expression optimization can be performed directly on the interface without switching tools or repeatedly retrieving conversation records, significantly improving review efficiency and content accuracy.
[0043] Once the proofreading and revisions are completed and the SOP knowledge points are confirmed to meet business standards, the release operation can be triggered with one click. Without the need for complicated approval processes, standardized SOP knowledge can be quickly accumulated into the enterprise knowledge base, ensuring that high-quality experience can be reused and disseminated in a timely manner.
[0044] As one implementation method, this method further includes the following steps: S7, Closed-loop optimization and continuous model iteration, specifically includes: Data Collection: The system automatically collects all modification records generated by experts during the review process. These records cover various operations such as content correction, logic adjustment, element supplementation, and expression optimization of AI-generated SOPs. Because the modifications are based on real business conversations and are completed by domain experts combining business standards and practical experience, each modification record has a high degree of business relevance and accuracy, constituting high-quality supervised learning samples and providing core data support for model iteration.
[0045] Model Fine-tuning: Based on the accumulated supervised learning sample library, the system will initiate the fine-tuning process of the large language model according to a preset cycle. By integrating the knowledge of business logic, expression norms, and element extraction experience contained in expert modifications into the model's training process, the model gradually learns and masters the SOP generation standards that meet the actual business needs, and continuously optimizes the accuracy of core element extraction, logical organization ability, and structured transformation efficiency.
[0046] The effect loop: By continuously operating a closed loop of "AI initially generates SOPs - expert review and feedback for modification - sample accumulation drives model fine-tuning - model generation quality improvement," the model's adaptability to business scenarios and its ability to grasp the core elements of SOPs will gradually improve. This process will directly manifest as a continuous improvement in the accuracy, completeness, and logical coherence of AI-generated SOPs, thereby significantly reducing the workload of manual review and modification, and even gradually achieving "AI generation - direct release" in some scenarios, forming a positive cycle of "model optimization - efficiency improvement," and continuously reducing business operating costs.
[0047] Example 2: This embodiment provides a SOP generation system based on a large language model, such as... Figure 2 As shown, it includes a data input and preprocessing module, an SOP generation engine, a human-machine collaborative review and release module, and a closed-loop optimization and model iteration module.
[0048] The data input and preprocessing module is used to capture dynamic conversation data from multiple sources in real time, including ongoing call audio streams and real-time text streams from scenarios such as online customer service systems and IM tools. For captured call audio streams, the module quickly and continuously converts them into text streams using real-time speech-to-text technology, ensuring synchronization between speech and text. For all text streams (including those converted from audio and directly captured text), systematic preprocessing operations are performed, covering dynamic noise filtering, sensitive information anonymization, colloquial expression normalization, and industry abbreviation expansion, to improve the purity and standardization of the text data. After preprocessing, the module calls a large language model to perform real-time intent recognition on the text stream, focusing on whether the conversation contains a complete pattern involving problem-solving or a clear operational instruction intent. Once a conversation that meets the criteria is detected, the SOP generation engine is immediately triggered to start the subsequent process.
[0049] When the SOP generation engine is triggered, the system immediately starts the large language model processing engine, focusing on in-depth analysis and information extraction of the complete context of the current conversation. During the analysis, the engine links historical interaction data with real-time dialogue content to accurately capture key SOP elements such as the core of the problem, key points of operation, decision basis, and execution constraints. It then incrementally updates the database, mapping the extracted information in real-time to preset structured fields—including dynamically filling in the problem description, atomically splitting and adding operation steps, and supplementing core entity information—to form a clear and structured statement. Subsequently, the engine systematically sorts and integrates all structured fields, streamlining the sequence of operation steps, clarifying logical dependencies, and supplementing the complete context, ultimately forming a standard operating procedure that conforms to business standards.
[0050] Afterwards, the human-machine collaborative review and release module generates the standard operating procedure and synchronizes it to the visual review interface, submitting it to frontline agents or domain experts for manual review. At this point, staff have a clear memory of the conversation details, problem background, and solutions, enabling them to efficiently proofread, modify, confirm, and supplement the AI-generated content, promptly correcting any potential information discrepancies, missing steps, or logical inconsistencies. During the review process, the system automatically records the specific differences in manual modifications, obtaining a modification difference record (Diff), forming a complete feedback trajectory. Once the review confirms that it meets business requirements, staff can complete the release with a single click, publishing the standardized operating procedure to the enterprise knowledge base for reuse and dissemination in subsequent business scenarios.
[0051] The closed-loop optimization and model iteration module automatically collects all modification records generated during the manual review process. These records, containing experts' business experience and standard specifications, serve as high-quality supervised learning samples. The system categorizes, organizes, and structures these samples to generate a dedicated fine-tuning dataset. It then periodically conducts supervised fine-tuning of the large language model based on this dataset, integrating experts' modification logic, expression standards, and element extraction experience into the model training process. Through precise feedback provided by human-machine collaborative review and a feedback-based closed-loop optimization mechanism, the professionalism and accuracy of the SOP are doubly guaranteed. Simultaneously, the model's scenario adaptability, information extraction accuracy, and structured transformation efficiency continuously improve, and the quality of AI-generated SOP drafts is constantly optimized, gradually reducing the workload and modification costs of manual review.
[0052] Example 3: This embodiment provides a computer-readable storage medium containing a computer program, which, when executed by one or more processors, causes the one or more processors to perform the SOP generation method based on a large language model provided in Embodiment 1 of this specification.
[0053] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0054] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0055] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0056] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0057] The implementation of all or part of the processes in the methods of the above embodiments can also be accomplished by a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the various method embodiments described above.
[0058] The embodiments described above are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for generating Standard Operating Procedures (SOPs) based on a large language model, characterized in that, Includes the following steps: S1. Acquire the audio or text stream of the session in real time, convert the audio stream into a text stream, and preprocess the text stream; S2. Perform real-time intent recognition on the preprocessed text stream using a large language model; S3. When an intention involving problem solving or operation instructions is identified, the context of the preprocessed text stream is analyzed and information is extracted. S4. Using incremental database updates, the analyzed and refined information is mapped in real time to preset structured fields to form structured statements. S5. Organize and integrate the structured fields to obtain the standard operating procedure.
2. The SOP generation method based on a large language model according to claim 1, characterized in that, The process of converting the call audio stream into a text stream includes: In multi-party calls, the system identifies the different speakers in real time and uses an automatic speech recognition model that has been fine-tuned on datasets that are specific to business domains and include common dialect accents for speech recognition and conversion.
3. The SOP generation method based on a large language model according to claim 1, characterized in that, In step S1, the preprocessing includes: Segment the continuous text stream into meaningful semantic segments or process it according to a fixed time window; Denoising and anonymizing of sensitive information are performed on the text stream; The text stream is normalized, spoken language is converted into written language, and terms and abbreviations are expanded and annotated.
4. The SOP generation method based on a large language model according to claim 1, characterized in that, The analysis and information extraction of the current session context includes: The core entities in the conversation are marked, and the subjects and objects of operation omitted in the dialogue are inferred based on these core entities, so as to extract key information.
5. The SOP generation method based on a large language model according to claim 1, characterized in that, The process of sorting and integrating the resulting structured fields to obtain the standard operating procedure includes: By integrating fragmented structured fields through a large language model, and obtaining the sequence of operation steps based on the timeline and logical dependencies, and supplementing the context, a standard operating procedure is obtained.
6. The SOP generation method based on a large language model according to claim 1, characterized in that, The method further includes: After obtaining the standard operating procedures, they are manually reviewed and then released.
7. The SOP generation method based on a large language model according to claim 6, characterized in that, The method further includes: The manually modified records generated from the manual review are collected as high-quality supervised learning samples. The large language model was fine-tuned using supervised learning samples.
8. A SOP generation system based on a large language model, characterized in that, Includes a data input and preprocessing module and a SOP generation engine: The data input and preprocessing module is used to capture the audio stream of an ongoing call or the text stream in real time, convert the audio stream into a text stream, preprocess the text stream, and perform intent recognition on the preprocessed text stream using a large language model to determine whether there is a conversation involving problem solving or operation instruction intent, and trigger the SOP generation engine. When the SOP generation engine is triggered, the large language model processing engine is started to analyze and extract information from the context of the current session. The extracted information is mapped to the preset structured fields in real time through incremental database updates to form structured statements. The structured fields are then sorted and integrated to obtain the standard operating procedure.
9. The SOP generation system based on a large language model according to claim 8, characterized in that, The system also includes a human-machine collaborative review and release module and a closed-loop optimization and model iteration module; The human-machine collaborative review and release module is used to obtain the standard operating procedure, submit the output standard operating procedure to human review, and release it after review. The closed-loop optimization and model iteration module is used to collect manually modified records as high-quality supervised learning samples, and to fine-tune the large language model using supervised learning samples.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by one or more processors, it causes the one or more processors to perform the steps of the method according to any one of claims 1-7.