Voice interaction safety knowledge question answering method and device based on large model, storage medium and electronic equipment

Through the voice interaction security knowledge question and answer method based on the big model, voice and image acquisition controls are used to process the big model in combination with security knowledge, the problems of inefficiency of traditional security management methods and complex interaction interfaces are solved, and convenient and efficient security knowledge interaction solutions are provided.

CN120296127APending Publication Date: 2025-07-11CHINA NAT BUILDING MATERIALS TECH CO LTD +3
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510367464.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional security management methods rely on manual operations, are inefficient and prone to errors. The interactive interface of smart assistants or security management platforms is complex, has poor operational convenience, and lacks accuracy in security knowledge interaction.

Method used

The voice interaction security knowledge question and answer method based on the big model is adopted to obtain interactive information through the voice acquisition control and image acquisition control, and the large model recognizes semantic information using pre-trained security knowledge, and match knowledge point information from the security knowledge database to provide security knowledge replies.

Benefits of technology

It realizes a simple operation interactive interface, expands the scope of application of the user group, and is especially suitable for older users, improves the efficiency and accuracy of security knowledge interaction, simplifies information input methods, and realizes quick question-and-answer and analysis of security knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296127A_ABST
    Figure CN120296127A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction safety knowledge question answering method and device based on a large model, a storage medium and electronic equipment. The method comprises the following steps: displaying an interaction page, wherein the interaction page comprises a voice acquisition control and an image acquisition control; obtaining interaction voice collected by a voice collection control, and / or obtaining at least one image frame collected by an image collection control; identifying target semantic information corresponding to the interactive voice and / or the at least one image frame through a pre-trained security knowledge processing large model, matching from a security knowledge database through the target semantic information to obtain at least one piece of knowledge point information, and obtaining security knowledge reply information based on the knowledge point information; and outputting the safety knowledge reply information through the interaction page. The process of manual analysis and safety knowledge extraction is replaced, rapid question answering and analysis of the safety knowledge are achieved, and the technical effects of convenience and high efficiency are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of work safety technologies, and in particular, to a voice interaction safety knowledge Q&A method, device, storage medium, and electronic device based on a large model. Background Art

[0002] With the rapid development of information technology, enterprises' requirements for safety management have become increasingly strict.

[0003] Traditional safety management methods rely heavily on manual operations. For example, tasks such as safety knowledge query, regulation interpretation, and analysis of business safety data are all completed manually. This method is extremely inefficient and prone to errors.

[0004] Currently, there are already some intelligent assistants and safety management platforms applied in various fields, which can replace the above-mentioned manual operations. There are at least the following technical problems in the prior art: the interaction interfaces of current intelligent assistants or safety management platforms are complex and the operation convenience is poor; and, the interaction accuracy of safety knowledge of current intelligent assistants or safety management platforms is poor. Summary of the Invention

[0005] The present disclosure provides a voice interaction safety knowledge Q&A method, device, storage medium, and electronic device based on a large model to provide a convenient interaction interface and improve the interaction efficiency and accuracy of safety knowledge.

[0006] According to one aspect of the present disclosure, there is provided a voice interaction safety knowledge Q&A method based on a large model, including:

[0007] Display an interaction page, where the interaction page includes a voice collection control and an image collection control;

[0008] Obtain the interaction voice collected through the voice collection control, and / or obtain at least one image frame collected through the image collection control;

[0009] Through a pre-trained safety knowledge processing large model, identify the target semantic information corresponding to the interaction voice and / or the at least one image frame, perform matching from a safety knowledge database through the target semantic information to obtain at least one knowledge point information, and obtain a safety knowledge reply information based on the knowledge point information;

[0010] Output the safety knowledge reply information through the interaction page.

[0011] Optionally, the safety knowledge processing large model includes a voice processing sub-model, an image processing sub-model, and a semantic integration sub-model;

[0012] Wherein, the semantic recognition of the interactive speech is performed by the speech processing sub-model to obtain first semantic information; and / or, the semantic recognition of the at least one image frame is performed by the image processing sub-model to obtain second semantic information;

[0013] The first semantic information and / or the second semantic information are semantically integrated by the semantic integration sub-model to obtain the target semantic information.

[0014] Optionally, the method further includes: displaying speech calibration text on the interactive page; obtaining the calibration speech corresponding to the speech calibration text, and extracting the accent feature and / or dialect feature corresponding to the calibration speech; during the process of the speech processing sub-model processing the interactive speech, inputting the accent feature and / or dialect feature into the speech processing sub-model as prior information.

[0015] Optionally, the method further includes: semantically integrating the at least one piece of knowledge point information to obtain an integrated text, and using the integrated text as the safety knowledge reply information.

[0016] Moreover, the safety knowledge processing large model further includes: an information integration sub-model configured to semantically integrate the at least one piece of knowledge point information to obtain an integrated text

[0017] Optionally, the at least one piece of knowledge point information includes one or more of safety knowledge clauses, safety knowledge cases, and safety operation templates;

[0018] The method further includes: obtaining a page jump voice, performing a page jump on the interactive page, and displaying the page after the jump, where the page after the jump displays one or more of the safety knowledge clauses, the safety knowledge cases, and the safety operation templates.

[0019] Optionally, the safety knowledge database includes knowledge point information corresponding to multiple business types;

[0020] The safety knowledge database is respectively communicatively connected to servers corresponding to multiple business types, obtains incremental knowledge point information of each of the multiple business types, and updates the safety knowledge database.

[0021] Optionally, the method further includes: displaying interactive prompt information on the interactive page, where the interactive prompt information is used to guide the user's next round of interaction.

[0022] According to another aspect of the present disclosure, there is provided a speech interaction safety knowledge answering device based on a large model, including:

[0023] A page display module for displaying an interactive page, where the interactive page includes a voice collection control and an image collection control;

[0024] A data collection module for obtaining the interactive voice collected through the voice collection control and / or obtaining at least one image frame collected through the image collection control;

[0025] A data processing module for identifying the target semantic information corresponding to the interactive voice and / or the at least one image frame through a pre-trained large model for security knowledge processing, matching from a security knowledge database through the target semantic information to obtain at least one knowledge point information, and obtaining a security knowledge reply information based on the knowledge point information;

[0026] A security knowledge reply information output module for outputting the security knowledge reply information through the interactive page.

[0027] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device includes:

[0028] At least one processor; and

[0029] A memory communicatively connected to the at least one processor; wherein,

[0030] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the method for answering security knowledge in voice interaction based on a large model according to any embodiment of the present disclosure.

[0031] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method for answering security knowledge in voice interaction based on a large model according to any embodiment of the present disclosure is implemented.

[0032] The technical solution of the embodiment of the present disclosure provides an interaction page including a voice acquisition control and an image acquisition control. This interaction page is concise, easy to operate, and applicable to a wide range of user groups. In particular, it provides a user-friendly operation page for elderly users. By using a large model for processing safety knowledge to perform semantic extraction on the acquired interaction voice and / or image frames, target semantic information is obtained, simplifying the information input method and reducing the user's usage threshold. By storing knowledge point information related to work safety in a safety knowledge database, and the knowledge point information is accurate and effective. At least one piece of knowledge point information is obtained by matching the target semantic information in the safety knowledge database, so as to obtain the safety knowledge reply information corresponding to the interaction voice and / or image frames and output it, replacing the process of manually analyzing and extracting safety knowledge, realizing rapid question answering and analysis of safety knowledge, and achieving the technical effect of convenience and high efficiency.

[0033] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 is a flowchart of a method for voice interaction safety knowledge question answering based on a large model provided by an embodiment of the present disclosure;

[0036] Figure 2 is a schematic diagram of the interaction page provided by an embodiment of the present disclosure;

[0037] Figure 3 is a schematic diagram of the processing process of interaction voice and image frames provided by an embodiment of the present disclosure;

[0038] Figure 4 is a schematic diagram of the jump of the interaction page provided by an embodiment of the present disclosure;

[0039] Figure 5 is a schematic diagram of the structure of a device for voice interaction safety knowledge question answering based on a large model provided by an embodiment of the present disclosure;

[0040] Figure 6 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To enable those skilled in the art to better understand the present disclosure solution, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0043] Figure 1 is a flowchart of a method for safe knowledge Q&A in voice interaction based on a large model provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of improving the operation convenience of users on the interaction page by displaying a simplified interaction page, realizing convenient information collection by collecting voice and / or image frames, and realizing efficient and accurate safe knowledge interaction by processing voice and / or image frames through a large model for safe knowledge processing. This method can be executed by a device for safe knowledge Q&A in voice interaction based on a large model, and the device for safe knowledge Q&A in voice interaction based on a large model can be implemented in the form of hardware and / or software, and the device for safe knowledge Q&A in voice interaction based on a large model can be configured in a computer and a mobile terminal. Among them, the mobile terminal may include, but is not limited to, mobile phones, tablet computers, etc. As Figure 1 shown, the method includes:

[0044] S110. Display an interaction page, where the interaction page includes a voice collection control and an image collection control.

[0045] In this embodiment, an interaction page is provided, and the interaction with the user is realized through the voice collection control and the image collection control in the interaction page. Exemplarily, see Figure 2 , Figure 2 is a schematic diagram of the interaction page provided by an embodiment of the present disclosure. It can be understood that the positions of the voice collection control and the image collection control in the interaction page are not limited.

[0046] S120. Obtain the interactive voice collected by the voice collection control, and / or obtain at least one image frame collected by the image collection control.

[0047] In response to the triggering operation on the voice collection control, start the pickup, and obtain the interactive voice collected during the triggering process of the voice collection control. In response to the triggering operation on the image collection control, start the camera, and obtain the image frames collected during the triggering process of the image collection control. The image frames can be one or more, and multiple image frames form a video.

[0048] The voice collection control and the image collection control on the interactive page are simple and convenient to operate, achieving the effect of simplifying the interactive page and reducing the operation difficulty, and providing a friendly operation experience for the interactive process related to safety knowledge management.

[0049] In some embodiments, use the interactive voice as the input information of the safety knowledge processing large model to obtain the safety knowledge reply information corresponding to the interactive voice. The interactive voice can include a query statement about safety knowledge. Correspondingly, the safety knowledge reply information is the query result corresponding to the query statement. Implement safety knowledge query and Q&A through the safety knowledge processing large model and the safety knowledge database, which is convenient, efficient, and improves the query efficiency. Exemplarily, the interactive voice can be "Query the safety operation process in the high-altitude operation scenario", "Is there any safety risk in using xx tool in the xx scenario", etc.

[0050] In some embodiments, use at least one image frame as the input information of the safety knowledge processing large model to obtain the safety knowledge reply information. At least one image frame can include the image frames collected for the object to be analyzed. For example, the object to be analyzed can be a business scenario. Taking the construction industry as an example, the business scenario can include, but is not limited to, high-altitude operation scenarios, hot work operation scenarios, earthwork operation scenarios, demolition operation scenarios, confined space operation scenarios, and temporary power use operation scenarios, etc. It can be understood that the business scenarios corresponding to different industries can be different. By collecting image frames of the business scenario, analyzing and processing the image frames of the business scenario through the safety knowledge processing large model and the safety knowledge database, the obtained safety knowledge reply information can include one or more of the safety knowledge corresponding to the business scenario, the safety risks existing in the business scenario, etc.

[0051] For example, the object to be analyzed can be business data. Correspondingly, the image frames can include the above-mentioned business data, and the business data can include one or more of safety event record data and equipment operation status data. By analyzing and processing the image frames of the business data through the safety knowledge processing large model and the safety knowledge database, the obtained safety knowledge reply information can include one or more of the safety knowledge corresponding to the business data, the safety risks existing in the business data, etc.

[0052] In some embodiments, interactive voice and at least one image frame are used as input information for a security knowledge processing large model to obtain security knowledge response information. The image frame includes an image frame collected for the object to be analyzed. For example, the object to be analyzed can be a business scenario or business data. The interactive voice can be understood as including a voice with processing requirements for at least one image frame. For example, the interactive voice can be "Identify security risks existing in the business scenario in the video", "Query the safe operation process of the business scenario in the image", "Analyze whether there are abnormalities in the business data in the image", etc. Through the security knowledge processing large model and the security knowledge database, the interactive voice and at least one image frame are processed to obtain security knowledge response information corresponding to the processing requirements of at least one image frame.

[0053] S130. Through a pre-trained security knowledge processing large model, identify the target semantic information corresponding to the interactive voice and / or the at least one image frame, match through the target semantic information in the security knowledge database to obtain at least one piece of knowledge point information, and obtain security knowledge response information based on the knowledge point information.

[0054] Based on the above embodiments, by combining the security knowledge processing large model and the security knowledge database, the interactive voice and / or at least one image frame are processed to implement processing processes such as question answering and analysis of security knowledge. Among them, the security knowledge processing large model is a machine learning model, such as a neural network model, which is used to extract target voice information from the interactive voice and / or image frame. The security knowledge database is a database for storing knowledge point information related to work safety. For example, the knowledge point information in the security knowledge database includes, but is not limited to, laws and regulations documents, industry standard specifications, enterprise rules and regulations, and case reports. The target voice information extracted based on the security knowledge processing large model is matched with the knowledge point information in the security knowledge database to obtain security knowledge response information.

[0055] Optionally, the security knowledge database includes knowledge point information corresponding to multiple business types; among them, the multiple business types here can be the business types involved by the enterprise. Taking an enterprise in the construction industry as an example, the business types can include, but are not limited to, high-altitude operation types, hot work operation types, earthwork operation types, demolition operation types, confined space operation types, and temporary power use operation types, etc. Different business types can correspond to different knowledge point information. By storing knowledge point information corresponding to multiple business types in the security knowledge database, the integration of security knowledge for different business types is realized, the data silos are broken, and the security knowledge of different business types is made to be interconnected and penetrate, improving the comprehensiveness of security knowledge.

[0056] The safety knowledge database is dynamically updated. By timely updating the safety knowledge database, the accuracy and effectiveness of the safety knowledge in the safety knowledge database are improved. Optionally, the safety knowledge database is communicatively connected to servers corresponding to multiple service types respectively, obtains incremental knowledge point information of each of the multiple service types, and updates the safety knowledge database. Herein, each service type may correspond to one or more servers, or one server corresponds to multiple service types, which is not limited herein.

[0057] In some embodiments, the safety knowledge database may be updated regularly. At the update time, data requests are sent to servers corresponding to multiple service types respectively to obtain incremental knowledge point information of each service type, and the safety knowledge database is updated based on the incremental knowledge point information of the service type. Herein, the update time interval of the safety knowledge database may be set according to the update requirement. In some embodiments, when incremental knowledge point information is detected in the server of any service type, an update instruction is generated. The update instruction carries the incremental knowledge point information, and the update instruction is transmitted to the safety knowledge database to update the safety knowledge database.

[0058] Updating the safety knowledge database may include adding the incremental knowledge point information to the safety knowledge database, or replacing the stored knowledge point information in the safety knowledge database based on the incremental knowledge point information. Optionally, the incremental knowledge point information includes a knowledge point name and a version number. The incremental knowledge point name is matched in the safety knowledge database. A failed match indicates that there is no knowledge point information corresponding to the knowledge point name in the safety knowledge database, and the incremental knowledge point information is added to the safety knowledge database. A successful match indicates that there is already knowledge point information corresponding to the knowledge point name in the safety knowledge database. The version number of the incremental knowledge point information and the version number of the stored knowledge point information that matches successfully are compared. If the version number of the incremental knowledge point information is greater than the version number of the stored knowledge point information, the stored knowledge point information that is matched is replaced based on the incremental knowledge point information.

[0059] It can be understood that laws and regulations documents, industry standard specifications, enterprise rules and regulations, and operation risk sources may change. Correspondingly, the knowledge point information related to work safety may be newly added or changed. By timely updating the knowledge point information in the safety knowledge database, problems such as expired knowledge point information and deviation of safety management focus caused by changes in laws and regulations documents, industry standard specifications, enterprise rules and regulations, and operation risk sources can be avoided, and various changes in the enterprise environment can be flexibly adapted to ensure the accuracy and effectiveness of the knowledge point information in the safety knowledge database, and further improve the accuracy and effectiveness of the safety knowledge reply information.

[0060] Based on the above embodiments, the security knowledge processing large model includes a speech processing sub-model, an image processing sub-model, and a semantic integration sub-model; wherein, the semantic recognition of the interactive speech is performed by the speech processing sub-model to obtain the first semantic information; and / or, the semantic recognition of the at least one image frame is performed by the image processing sub-model to obtain the second semantic information; the first semantic information and / or the second semantic information are semantically integrated by the semantic integration sub-model to obtain the target semantic information.

[0061] The speech processing sub-model can be, for example, a transformer model, which is used to extract semantic information from the interactive speech to obtain the first semantic information; the specific structure and type of the speech processing sub-model are not limited here. The image processing sub-model can be, for example, a transformer model. Exemplarily, the image processing sub-model is a Vision Transformer (ViT) model, which is used to extract semantics from at least one image frame to obtain the second semantic information; the specific structure and type of the image processing sub-model are not limited here. The semantic integration sub-model can be, for example, a convolutional sub-model, which is used to integrate the first semantic information and the second semantic information to obtain the target semantic information.

[0062] The first semantic information, the second semantic information, and the target semantic information here can be in vector form or text form. In the case of only processing the interactive speech, the first semantic information is used as the target semantic information; in the case of only processing at least one image frame, the second semantic information is used as the target semantic information. In the case of processing the interactive speech and at least one image frame, the first semantic information can be text (or vector) representing the processing requirement; the second semantic information can be the description text (or vector) of at least one image frame. Correspondingly, the target semantic information is the text (or vector) of the processing requirement for the description text of at least one image frame.

[0063] By matching the target semantic information in the security knowledge database, the knowledge point information that matches the target semantic information is determined. Specifically, the similarity data between the target semantic information and the knowledge point information in the security knowledge database is determined respectively. This similarity data can be calculated by a preset similarity algorithm, and the preset similarity algorithm includes but is not limited to the cosine similarity algorithm. The knowledge point information in the security knowledge database is sorted based on the similarity data, and at least one knowledge point information that matches successfully is determined from the sorted knowledge point information based on the similarity threshold and / or the knowledge point quantity threshold.

[0064] Optionally, at least one knowledge point information that matches successfully is used as the security knowledge reply information.

[0065] Optionally, semantic integration is performed on the at least one knowledge point information to obtain an integrated text, and the integrated text is used as the security knowledge reply information. Through semantic integration, multiple independent knowledge point information is integrated into a semantically fluent security knowledge reply information, improving the integrity and fluency of the security knowledge reply information and avoiding the problem of fragmented results.

[0066] The security knowledge processing large model further includes an information integration sub-model, and the information integration sub-model is configured to perform semantic integration on the at least one knowledge point information to obtain an integrated text. For example, the information integration sub-model can be a convolutional sub-model; for example, the information integration sub-model can be a bert model. Here, the specific structure and type of the information integration sub-model are not limited.

[0067] Figure 3 It is a schematic diagram of the processing process of interactive voice and image frames provided by an embodiment of the present disclosure; through Figure 2 The interactive page shown collects at least one image frame of the interactive voice including the processing requirement and the object to be analyzed. The semantic extraction of the interactive voice is performed through the voice processing sub-model to obtain the first semantic information, and the semantic extraction of the at least one image frame is performed through the image processing sub-model to obtain the second semantic information. The semantic integration sub-model performs semantic integration on the first semantic information and the second semantic information to obtain the target semantic information, matches the target semantic information in the security knowledge database to obtain at least one knowledge point information, and inputs the at least one knowledge point information into the information integration sub-model to obtain the security knowledge reply information. Among them, the information integration sub-model can be connected to the semantic integration sub-model, and the semantic integration sub-model inputs the target semantic information into the information integration sub-model, which serves as reference information for the information integration sub-model to integrate the at least one knowledge point information, so that the integrated security knowledge reply information matches the target semantic information and improves the accuracy of information integration.

[0068] Based on the above embodiments, it can be understood that the interactive voice can be Mandarin or a dialect, and the accents of different users can also be different. To improve the accuracy of interactive voice recognition, the method further includes: displaying voice calibration text on the interactive page; obtaining the calibration voice corresponding to the voice calibration text, and extracting the accent features and / or dialect features corresponding to the calibration voice; during the process of the voice processing sub-model processing the interactive voice, inputting the accent features and / or dialect features into the voice processing sub-model as prior information.

[0069] The voice calibration text can be used to calibrate the user's dialect and accent. By displaying the voice calibration text on the interaction page, the calibrated voice collected when the user reads the voice calibration text is obtained. Optionally, through a pre-set acoustic feature extraction module, by inputting the calibrated voice into the acoustic feature extraction module, the accent feature and / or dialect feature corresponding to the calibrated voice is obtained. Optionally, the calibrated voice is transmitted to the server so that the server extracts the accent feature and / or dialect feature corresponding to the calibrated voice through the acoustic feature extraction module.

[0070] When the voice processing sub-model extracts the first semantic information of the interaction voice, it includes: converting the interaction voice into interaction text and extracting the first semantic information of the interaction text. In the process of converting the interaction voice into interaction text, using the accent feature and / or dialect feature as prior information is beneficial to improving the accuracy of the voice-to-text conversion, avoiding the problem of inaccurate recognition of the interaction semantics when the interaction voice is non-Mandarin and / or has an accent, further improving the accuracy of the target semantic information and the accuracy of the safety knowledge reply information.

[0071] In some embodiments, the recognized interaction text is displayed on the interaction page so that the user can judge the accuracy of the interaction text and can re-enter the voice in a timely manner if the recognition is accurate.

[0072] It can be understood that each user corresponds to an account. When each account logs in for the first time, the accent feature and / or dialect feature of the user is obtained and stored, so as to provide assistance for the voice processing sub-model to process the interaction voice in the subsequent processing process.

[0073] S140. Output the safety knowledge reply information through the interaction page.

[0074] The output method of the safety knowledge reply information can include: displaying the safety knowledge reply information on the interaction page and / or playing the voice corresponding to the safety knowledge reply information.

[0075] Optionally, playing the voice corresponding to the safety knowledge reply information includes converting the safety knowledge reply information into a voice with an accent feature and / or dialect feature based on the accent feature and / or dialect feature and playing it, which is convenient for the user to receive the safety knowledge reply information.

[0076] The technical solution provided in this embodiment provides an interaction page including a voice collection control and an image collection control. This interaction page is simple, convenient to operate, and applicable to a wide range of user groups. In particular, it provides a friendly operation page for elderly users. By using a large model for processing safety knowledge to perform semantic extraction on the collected interaction voice and / or image frames, target semantic information is obtained, simplifying the information input method and reducing the user's usage threshold. By storing knowledge point information related to safe production in a safety knowledge database, and the knowledge point information is accurate and effective. At least one knowledge point information is obtained by matching the target semantic information in the safety knowledge database, so as to obtain the safety knowledge reply information corresponding to the interaction voice and / or image frames and output it, replacing the process of manually analyzing and extracting safety knowledge, realizing the rapid question answering and analysis of safety knowledge, and achieving the technical effect of convenience and high efficiency.

[0077] On the basis of the above embodiment, the safety knowledge database may include legal and regulatory documents, industry standard specifications, enterprise rules and regulations, historical safety knowledge cases, and safety operation templates corresponding to each business type. Correspondingly, the at least one knowledge point information includes one or more of safety knowledge clauses, safety knowledge cases, and safety operation templates; among them, the safety knowledge clauses can be understood as knowledge point information matched from legal and regulatory documents, industry standard specifications, enterprise rules and regulations, etc.; the safety knowledge cases can be matched from historical safety knowledge cases, and the safety operation templates can be matched from the safety operation templates corresponding to each business type.

[0078] On the basis of displaying the safety knowledge reply information on the interaction page, it is also possible to display the basis for generating the safety knowledge reply information, that is, at least one knowledge point information.

[0079] Optionally, the method further includes: obtaining a page jump voice, performing a page jump on the interaction page, and displaying the page after the jump. One or more of the safety knowledge clauses, the safety knowledge cases, and the safety operation templates are displayed on the page after the jump. Among them, for example, the page jump voice can be "display safety knowledge clauses", or "jump to the display page of safety knowledge clauses", etc.

[0080] Exemplarily, refer to Figure 4 , Figure 4 is a schematic diagram of a page jump provided by an embodiment of the present disclosure. Figure 4 The left figure of is a schematic diagram of the interaction page. The interaction text corresponding to the interaction voice, the image frame, and the safety knowledge reply information are displayed on this interaction page. After the page jump voice is collected, the interaction page is controlled to jump to the display page of the safety knowledge clauses, and the safety knowledge clauses corresponding to the safety knowledge reply information are displayed, that is Figure 4 as shown in the right figure of.

[0081] On this basis,Figure 4 The left figure also includes controls corresponding to each type of knowledge point information. In response to a triggering operation on any control, the interactive page is switched to the display page corresponding to the control, and the knowledge point information of the type corresponding to the control is displayed. For example, clicking on the clause control will cause the interactive page to jump to the display page of the safety knowledge clause, and display the safety knowledge clause corresponding to the safety knowledge reply information.

[0082] In this embodiment, by inputting the page jump voice, the jump of the interactive page is realized, which simplifies the operation method of page jump and improves the operation convenience.

[0083] Based on the above embodiment, the method further includes: displaying interactive prompt information on the interactive page, where the interactive prompt information is used to guide the user's next round of interaction. The user can input the next round of interactive voice according to the displayed interactive prompt information. Among them, the interactive prompt information output by the safety knowledge processing large model can be one or more items.

[0084] Optionally, the interactive prompt information can be generated by the safety knowledge processing large model.

[0085] Optionally, the interactive prompt information can be determined according to the user's multi-round interactive voice in the historical interaction process. The determination includes the historical interaction process of the interactive voice in the current interaction process, determines the occurrence frequency of the next round of interactive voice of the interactive voice in the historical interaction process, and determines one or more interactive prompt information based on the occurrence frequency of the next round of interactive voice. It can be to sort the next round of interactive voice based on the occurrence frequency, and use the first n next round of interactive voice as the interactive prompt information.

[0086] By displaying the interactive prompt information, it provides prompt information for the next round of interaction, simplifies the interaction process, and reduces the threshold of the interaction process.

[0087] Based on the above embodiment, the embodiments of the present disclosure also provide an example of a method for processing the security of voice interaction data of a large model. Deploy the safety knowledge processing large model in an electronic device and train the safety knowledge processing large model. Among them, the voice processing sub-model, image processing sub-model, semantic integration sub-model, and information integration sub-model in the safety knowledge processing large model can be trained independently or as a whole to obtain a primary safety knowledge processing large model. The parameters of the primary safety knowledge processing large model are adjusted through a sample data set in a specific field to obtain a trained safety knowledge processing large model. The specific field can be the field to which the enterprise belongs, for example, it can be the construction field.

[0088] Build a safety knowledge database. The safety knowledge database can be deployed locally on an electronic device or on a server. The electronic device is communicatively connected to the server and can access the safety knowledge database. It can be understood that the safety knowledge database is continuously updated to ensure the accuracy and effectiveness of the knowledge point information in the safety knowledge database. Optimize and train the pre-trained large model for safety knowledge processing regularly to ensure the performance of the large model for safety knowledge processing, improve the accuracy of semantic extraction for interactive voice and image frames, and improve the accuracy of safety knowledge response information.

[0089] In one example, through Figure 2 the interactive page shown, obtain the interactive voice. The interactive voice can be a safety knowledge query voice. Through the voice processing sub-model in the large model for safety knowledge processing, identify the first semantic information corresponding to the interactive voice, use the first semantic information as the target semantic information, match at least one knowledge point information in the safety knowledge database, integrate the at least one knowledge point information through the information integration sub-model to obtain the safety knowledge response information, display the safety knowledge response information on the interactive page, and play the voice corresponding to the safety knowledge response information. Exemplarily, the safety knowledge query voice can be a query voice for the safety operation process of a certain business scenario, and the corresponding safety knowledge response information is the safety operation process of this business scenario.

[0090] In one example, through Figure 2 the interactive page shown, obtain the interactive voice and at least one image frame of a certain business scenario (the object to be analyzed, for example, it can be a hot work scenario). The interactive voice can be a voice containing the processing requirement for at least one image frame of the business scenario, and the processing requirement can be the safety risk analysis of the business scenario. Through the voice processing sub-model in the large model for safety knowledge processing, identify the first semantic information corresponding to the interactive voice, through the image and voice processing sub-model, identify the second semantic information corresponding to the at least one image frame, integrate the first semantic information and the second semantic information into the target semantic information through the semantic integration sub-model, match at least one knowledge point information in the safety knowledge database, integrate the at least one knowledge point information through the information integration sub-model to obtain the safety knowledge response information, display the safety knowledge response information on the interactive page, and play the voice corresponding to the safety knowledge response information. The safety knowledge response information can be the identification result of the safety risks existing in the business scenario, for example, including one or more of the safety risk types and risk levels.

[0091] Figure 5 is a schematic structural diagram of a voice interaction safety knowledge Q&A device based on a large model provided by an embodiment of the present disclosure. As Figure 5 shown, the device includes:

[0092] A page display module 210 for displaying an interactive page, where the interactive page includes a voice collection control and an image collection control;

[0093] A data collection module 220 for obtaining the interactive voice collected through the voice collection control and / or obtaining at least one image frame collected through the image collection control;

[0094] A data processing module 230 for identifying the target semantic information corresponding to the interactive voice and / or the at least one image frame through a pre-trained large safety knowledge processing model, matching through the target semantic information in a safety knowledge database to obtain at least one knowledge point information, and obtaining a safety knowledge reply information based on the knowledge point information;

[0095] A safety knowledge reply information output module 240 for outputting the safety knowledge reply information through the interactive page.

[0096] The technical solution of this embodiment provides an interactive page including a voice collection control and an image collection control. The interactive page is simple, convenient to operate, and has a wide range of applicable user groups. Especially, it provides a friendly operation page for elderly users. By using a large safety knowledge processing model to extract the semantics of the collected interactive voice and / or image frames to obtain the target semantic information, the information input method is simplified and the user usage threshold is reduced. By storing the knowledge point information related to work safety in a safety knowledge database, and the knowledge point information is accurate and effective. By matching the target semantic information in the safety knowledge database to obtain at least one knowledge point information, to obtain the safety knowledge reply information corresponding to the interactive voice and / or image frame, and output it, replacing the process of manually analyzing and extracting safety knowledge, realizing the rapid question answering and analysis of safety knowledge, and achieving the technical effect of convenience and high efficiency.

[0097] Based on the above embodiment, optionally, the large safety knowledge processing model includes a voice processing sub-model, an image processing sub-model, and a semantic integration sub-model;

[0098] The data processing module 230 is used to perform semantic recognition on the interactive voice through the voice processing sub-model to obtain the first semantic information; and / or perform semantic recognition on the at least one image frame through the image processing sub-model to obtain the second semantic information; and perform semantic integration on the first semantic information and / or the second semantic information through the semantic integration sub-model to obtain the target semantic information.

[0099] Optionally, the data processing module 230 is further configured to display the voice calibration text on the interaction page; obtain the calibration voice corresponding to the voice calibration text, and extract the accent feature and / or dialect feature corresponding to the calibration voice; during the process of the voice processing sub-model processing the interaction voice, input the accent feature and / or dialect feature into the voice processing sub-model as prior information.

[0100] Optionally, the data processing module 230 is further configured to perform semantic integration on the at least one piece of knowledge point information to obtain an integrated text, and use the integrated text as the safety knowledge reply information;

[0101] Moreover, the safety knowledge processing large model further includes: an information integration sub-model, which is configured to perform semantic integration on the at least one piece of knowledge point information to obtain an integrated text.

[0102] Based on the above embodiments, optionally, the at least one piece of knowledge point information includes one or more of safety knowledge clauses, safety knowledge cases, and safety operation templates;

[0103] The device further includes: a page processing module, configured to obtain a page jump voice, perform a page jump on the interaction page, and display the page after the jump, where the page after the jump displays one or more of the safety knowledge clauses, the safety knowledge cases, and the safety operation templates.

[0104] Based on the above embodiments, optionally, the safety knowledge database includes knowledge point information corresponding to multiple service types; the safety knowledge database is respectively communicatively connected to servers corresponding to multiple service types, obtains incremental knowledge point information of each of the multiple service types, and updates the safety knowledge database.

[0105] Based on the above embodiments, optionally, the safety knowledge reply information output module 240 is further configured to display interaction prompt information on the interaction page, and the interaction prompt information is used to guide the user's next round of interaction.

[0106] The voice interaction safety knowledge Q&A device based on a large model provided by the embodiments of the present disclosure can execute the voice interaction safety knowledge Q&A method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0107] Figure 6It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. The electronic device 10 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0108] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the random access memory (RAM) 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the read-only memory (ROM) 12, and the random access memory (RAM) 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0109] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0110] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the voice interaction security knowledge Q&A method based on a large model.

[0111] In some embodiments, the large model-based voice interaction security knowledge Q&A method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the read-only memory (ROM) 12 and / or the communication unit 19. When the computer program is loaded into the random access memory (RAM) 13 and executed by the processor 11, one or more steps of the large model-based voice interaction security knowledge Q&A method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the large model-based voice interaction security knowledge Q&A method by any other suitable means (e.g., by means of firmware).

[0112] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0113] The computer program for implementing the large model-based voice interaction security knowledge Q&A method of the present disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0114] The embodiments of the present disclosure also provide a computer-readable storage medium storing computer instructions for causing a processor to execute a large model-based voice interaction security knowledge Q&A method, the method including:

[0115] Display an interactive page, where the interactive page includes a voice collection control and an image collection control; obtain the interactive voice collected by the voice collection control, and / or obtain at least one image frame collected by the image collection control; through a pre-trained large model for processing security knowledge, identify the target semantic information corresponding to the interactive voice and / or the at least one image frame, match from a security knowledge database through the target semantic information to obtain at least one piece of knowledge information, and obtain a security knowledge reply information based on the knowledge information; output the security knowledge reply information through the interactive page.

[0116] In the context of the present disclosure, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0118] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend, middleware, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0119] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0120] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0121] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A voice interaction security knowledge Q&A method based on a large model, characterized in that, Including: Display an interactive page, where the interactive page includes a voice collection control and an image collection control; Obtain the interactive voice collected by the voice collection control, and / or obtain at least one image frame collected by the image collection control; Through a pre-trained large security knowledge processing model, identify the target semantic information corresponding to the interactive voice and / or the at least one image frame, match from the security knowledge database through the target semantic information to obtain at least one knowledge point information, and obtain security knowledge reply information based on the knowledge point information; Output the security knowledge reply information through the interactive page.

2. The method according to claim 1, characterized in that, The security knowledge processing large model includes a voice processing sub-model, an image processing sub-model, and a semantic integration sub-model; Among them, through the voice processing sub-model, semantic recognition is performed on the interactive voice to obtain first semantic information; and / or, through the image processing sub-model, semantic recognition is performed on the at least one image frame to obtain second semantic information; Through the semantic integration sub-model, the first semantic information and / or the second semantic information are semantically integrated to obtain the target semantic information.

3. The method according to claim 2, wherein The method further includes: Display voice calibration text on the interactive page; Obtain the calibration voice corresponding to the voice calibration text, and extract the accent feature and / or dialect feature corresponding to the calibration voice; During the process of the voice processing sub-model processing the interactive voice, input the accent feature and / or dialect feature as prior information into the voice processing sub-model.

4. The method according to claim 2, wherein The method further includes: Semantically integrate the at least one knowledge point information to obtain an integrated text, and use the integrated text as the security knowledge reply information; And, the security knowledge processing large model further includes: an information integration sub-model, and the information integration sub-model is configured to semantically integrate the at least one knowledge point information to obtain an integrated text.

5. The method according to claim 1, wherein The at least one knowledge point information includes one or more of security knowledge clauses, security knowledge cases, and security operation templates; The method further includes: Obtain a page jump voice, perform a page jump on the interactive page, and display the jumped page, where the jumped page displays one or more of the security knowledge clauses, the security knowledge cases, and the security operation templates.

6. The method according to claim 1, wherein The security knowledge database includes knowledge point information corresponding to multiple business types; The security knowledge database is respectively communicatively connected to servers corresponding to multiple business types, obtains incremental knowledge point information of each of the multiple business types, and updates the security knowledge database.

7. The method according to claim 1, characterized in that, The method further includes: displaying interactive prompt information on the interactive page, where the interactive prompt information is used to guide the user's next round of interaction.

8. A voice interaction security knowledge Q&A device based on a large model, characterized in that, Including: A page display module, configured to display an interactive page, where the interactive page includes a voice collection control and an image collection control; A data collection module, configured to obtain the interactive voice collected by the voice collection control, and / or obtain at least one image frame collected by the image collection control; A data processing module, configured to identify target semantic information corresponding to the interactive speech and / or the at least one image frame through a pre-trained large model for processing security knowledge, match from a security knowledge database through the target semantic information to obtain at least one piece of knowledge point information, and obtain security knowledge reply information based on the knowledge point information; A security knowledge reply information output module, configured to output the security knowledge reply information through the interactive page.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the large model-based voice interaction security knowledge Q&A method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the large model-based voice interaction security knowledge Q&A method according to any one of claims 1-7 when executed by a processor.

Citation Information

Patent Citations

  • Method and device for voice recognition

    CN105096940A

  • Equipment maintenance knowledge determination method and device, equipment and storage medium

    CN117764551A

  • Power transformation professional information intelligent interaction method

    CN119441438A

  • Intelligent question and answer method and system in building construction field

    CN119474328A

  • Question answering method and apparatus, and device and storage medium

    WO2024227415A1