Confirmation method and device of call key point, equipment, storage medium and program product

By constructing a conversation sequence for multi-person voice calls, and combining the call scenario and participant roles, the key points of the call are identified, solving the problem of incomplete capture of key points in multi-person voice calls, and achieving efficient and accurate identification of key points and immediate action.

CN120932635AActive Publication Date: 2025-11-11SHENZHEN JUBAOJIA INTELLIGENT TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511462140.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-11-11
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

In multi-person voice calls, existing technologies struggle to efficiently and accurately capture and confirm the key points of the call, and are prone to missing voice content other than that of the first speaker.

Method used

By acquiring the audio content of multi-person voice calls and converting it into text, pre-defined keywords for each participant are determined, a conversation sequence is constructed, and the key points of the call are identified by combining the call scenario and roles. NLP and semantic understanding technologies are then used to identify the key points of multi-person voice calls.

Benefits of technology

It enables more efficient and comprehensive identification of call priorities in multi-person voice calls, improves accuracy, and allows for timely implementation of necessary measures to adapt to different safety production scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932635A_ABST
    Figure CN120932635A_ABST
Patent Text Reader

Abstract

The invention discloses a call key point confirmation method and device, equipment, a storage medium and a program product, and relates to the technical field of voice interaction. The method comprises the following steps: acquiring a call scene of a multi-person voice call and voice content of each voice call participant, and converting the voice content into text content; determining a preset keyword of each voice call participant in the call scene according to the text content; a session sequence of the multi-person voice call is constructed, and the session sequence comprises a call scene, roles of all voice call participants in the call scene, and preset keywords of all the voice call participants in the call scene; and through the conversation sequence of the multi-person voice conversation, the conversation key point of the multi-person voice conversation is identified and obtained. The technical problem that call key points are not comprehensively, efficiently and accurately captured and confirmed in multi-person voice calls is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of voice interaction, and in particular to a method, device, equipment, storage medium, and computer program product for confirming call priorities. Background Technology

[0002] In current safety production scenarios, it is often necessary for participants in voice calls to wear smart safety helmets to conduct multi-person voice calls. When confirming the key points of a multi-person voice call, it is often necessary to first identify and determine the priority speaker among the participants, and then directly use the priority speaker's voice content as the key point of the call. This method is prone to missing the voice content of other participants besides the priority speaker, resulting in incomplete, inefficient and inaccurate capture and confirmation of the key points of a multi-person voice call.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a method, device, equipment, storage medium, and computer program product for confirming the key points of a call, aiming to solve the technical problems of incomplete, inefficient, and inaccurate capture and confirmation of key points in multi-person voice calls.

[0005] To achieve the above objectives, this application proposes a method for confirming the focus of a call, the method comprising: The system acquires the call scenario of a multi-person voice call and the voice content of each participant, and converts the voice content into text content. Based on the text content, determine the preset keywords for each voice call participant in the call scenario; Construct a conversation sequence for a multi-person voice call, wherein the conversation sequence includes the call scenario, the roles of each voice call participant in the call scenario, and preset keywords for each voice call participant in the call scenario; By analyzing the conversation sequence of a multi-person voice call, the key points of the call can be identified.

[0006] In one embodiment, the step of constructing a conversation sequence for a multi-person voice call includes: Obtain the number of participants in a multi-person voice call, and determine the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of participants. Based on the sequence window size of the session sequence to be constructed, a multi-person voice call session sequence is constructed.

[0007] In one embodiment, the step of determining the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of people in the voice call includes: A prediction model for the sequence window size is constructed using historical data, wherein the historical data includes historical call scenarios, historical voice call participants, and the historical sequence window size of the historical conversation sequences corresponding to the historical call scenarios and the historical voice call participants. The prediction model determines the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of people in the voice call.

[0008] In one embodiment, the step of constructing a multi-person voice call session sequence based on the sequence window size of the session sequence to be constructed includes: Obtain key historical calls within a historical time range and determine the risk level of those key historical calls; Adjust the sequence window size of the session sequence according to the risk level to obtain the latest sequence window size of the session sequence; Based on the latest sequence window size of the session sequence to be constructed, a multi-person voice call session sequence is constructed.

[0009] In one embodiment, after the step of determining the preset keywords of each voice call participant in the call scenario based on the text content, the method for confirming the call focus further includes: Each participant in the voice call is assigned to a task group, and the task scenario for each task group is obtained. Construct a conversation sequence for each task group, wherein the conversation sequence for each task group includes the task scenario, the roles of each voice call participant in the task scenario, and the preset keywords of each voice call participant in the task scenario. By analyzing the conversation sequences of task groups, the key points of each task group's calls can be identified. Based on the task scenarios and call priorities of each task group, determine the key points of multi-person voice calls.

[0010] In one embodiment, after the step of determining the call focus of a multi-person voice call based on the task scenario and call focus of each task group, the method further includes: The key points of a multi-person voice call, identified through the conversation sequence of the multi-person voice call, are taken as the first key point of the call. The call priorities for multi-person voice calls, determined based on the task scenarios and call priorities of each task group, will be used as the second call priorities. A third call focus is determined based on the first and second call focus, and this third call focus is used as the target call focus for multi-person voice calls.

[0011] Furthermore, to achieve the above objectives, this application also proposes a device for confirming call priorities, the device comprising: The conversion module is used to acquire the call scenario of a multi-person voice call and the voice content of each participant in the voice call, and convert the voice content into text content. The determination module is used to determine the preset keywords of each voice call participant in the call scenario based on the text content. A construction module is used to construct a conversation sequence for a multi-person voice call, wherein the conversation sequence includes the call scenario, the roles of each voice call participant in the call scenario, and the preset keywords of each voice call participant in the call scenario; The recognition module is used to identify the key points of a multi-person voice call by analyzing the conversation sequence of the multi-person voice call.

[0012] In addition, to achieve the above objectives, this application also proposes a device for confirming call priorities, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the call priorities confirmation method as described above.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the call focus confirmation method as described above.

[0014] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the call focus confirmation method described above.

[0015] One or more technical solutions proposed in this application have at least the following technical effects: In this application, based on the text content corresponding to the voice content of each participant in a multi-person voice call, preset keywords for each participant in the call scenario are determined. A conversation sequence is then constructed, consisting of the call scenario, the roles of each participant in that scenario, and the aforementioned preset keywords. This conversation sequence allows for the identification of the key points of the multi-person voice call. Therefore, even when participants are wearing smart helmets during a multi-person voice call, the key points can still be determined from the voice recordings of each participant, even without identifying and determining the first speaker or using their voice content as the key point. Because the identification and confirmation process of the first speaker is skipped, and the key points are directly determined from the voice recordings of each participant, and because the conversation sequence fully considers the call scenario, the roles of each participant in that scenario, and the preset keywords in their voice content, the identified key points are more efficient, comprehensive, and accurate. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating the first embodiment of the method for confirming call focus in this application; Figure 2 An application diagram illustrating the first embodiment of the call focus confirmation method of this application; Figure 3 A flowchart illustrating the second embodiment of the method for confirming call focus in this application; Figure 4 This is a schematic diagram of the module structure of the call focus confirmation device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the call focus confirmation method in this application embodiment.

[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0022] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, or a call focus confirmation device such as a voice call server. The following description uses a call focus confirmation device as an example to illustrate this embodiment and the subsequent embodiments.

[0023] Based on this, embodiments of this application provide a method for confirming the focus of a call, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the method for confirming the focus of a call in this application.

[0024] In this embodiment, the method for confirming the focus of the call includes steps S10 to S40: Step S10: Obtain the call scenario of the multi-person voice call and the voice content of each voice call participant, and convert the voice content into text content; Multi-person voice call scenarios refer to real-world safety production scenarios where multiple people wearing smart safety helmets engage in voice calls via the helmets. The voice content of each participant can be captured using a voice acquisition device, such as a microphone, integrated into the smart safety helmet or used in conjunction with it. In one embodiment, ASR (Automatic Speech Recognition) technology can be used to convert the voice content into text content.

[0025] Step S20: Based on the text content, determine the preset keywords for each voice call participant in the call scenario; In one embodiment, a set of high-frequency key instructions such as pause work, evacuate, report location, request support, and tool failure are predefined for different actual safety production scenarios (construction, inspection, emergency rescue, etc.), and custom thesaurus such as proprietary equipment names, terms, and specific instructions are allowed to be added.

[0026] Within the text content, specific types of entity information are identified, including: 1. Personnel / Positions, such as Zhang San, Wang Gong, and Safety Officer, used to confirm the object of the instruction; 2. Equipment / Tools, such as Crane and Welding Machine, used to confirm the operation target; 3. Location / Area, such as Platform 3, West Passage, and North Side of the Excavation Pit, used to confirm location information; 4. Actions / Instructions, such as Start, Stop, Raise, Lower, Check, and Repair, used to confirm the operation content; 5. Status / Problem, such as Overheating, Loosening, Leaking, and Breaking, used to describe the event; and 6. Time / Quantity, such as 10 minutes later, 30 meters, and 3 times, used to confirm key parameters. Furthermore, the aforementioned entity information can be used to confirm the preset keywords of each participant in the voice call scenario.

[0027] It should be noted that the preset keywords for each participant in a voice call may differ depending on the call scenario.

[0028] Step S30: Construct a conversation sequence for a multi-person voice call, wherein the conversation sequence includes the call scenario, the roles of each voice call participant in the call scenario, and the preset keywords of each voice call participant in the call scenario. The constructed multi-person voice call conversation sequence includes the call scenario, the roles of each voice call participant in the call scenario, and the preset keywords of each voice call participant in the call scenario. For example, a certain conversation sequence is a construction scenario. The roles of each voice call participant in the construction scenario are such as project manager / commander-in-chief, site engineer / technical head, safety supervisor, team leader / foreman, equipment operator (crane, excavator driver, etc.), worker / construction worker, logistics / materials manager, etc. The preset keywords of each voice call participant in the call scenario are such as commander-in-chief: start now...when will the crane arrive...each team reports the completion status of the node, where the keywords are start, crane, team report, node completion status; safety supervisor: clear personnel under the boom...wear safety belt for working at height, where the keywords are under the boom, clear personnel, working at height, safety belt.

[0029] Step S40: Identify the key points of the multi-person voice call through the conversation sequence of the multi-person voice call.

[0030] In one embodiment, NLP (Natural Language Processing) technology, along with semantic understanding and intent recognition, can be used to determine the call focus in a multi-person voice call sequence. The call focus refers to the core key information after summarizing the voice focus of each voice call participant.

[0031] Furthermore, after automatically identifying the key points of a call, the system immediately takes the necessary measures corresponding to those points. For example, if the key point of a call is identified as a construction site hazard with personnel injuries, the system automatically implements the corresponding preset safety plan, such as issuing alarms and emergency shutdown, thereby implementing the preset safety plan more quickly. For example, risk levels or importance levels can be set based on the key points of the call, such as high priority: trigger words (danger, leakage, collapse, call 120), safety procedure keywords, emergency instructions; medium priority: operation instructions, status reports, location updates; low priority: idle chat, background noise, non-critical descriptions. Based on the risk level or importance level, corresponding alarms or information highlighting are triggered.

[0032] Above, refer to Figure 2 When participants in a multi-person voice call are wearing smart helmets, even without identifying and determining the first speaker among all participants and using their voice content as the focus of the call, the focus of the call can still be determined from the voice recordings of all participants. Because it skips the identification and confirmation process of the first speaker and directly determines the focus from the voice recordings of all participants, and because the conversation sequence of the multi-person voice call fully considers the call scenario, the roles of each participant within that scenario, and the preset keywords in each participant's voice content, the identified focus of the multi-person voice call is more efficient, comprehensive, and accurate.

[0033] Using keywords from the audio content in the conversation sequence, rather than the text corresponding to the entire audio content, makes identifying the call focus based on the conversation sequence more efficient. Furthermore, since non-keywords other than keywords are removed, identifying the call focus based on the conversation sequence is also more accurate. In addition to using keywords from the audio content, the conversation sequence also incorporates the scenario of multi-person voice calls and the roles of each participant, making the identification of the call focus based on the conversation sequence more comprehensive and accurate.

[0034] In one feasible implementation, step S30 may include steps A11-A12: Step A11: Obtain the number of participants in the multi-person voice call, and determine the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of participants. Step A12: Based on the sequence window size of the session sequence to be constructed, a multi-person voice call session sequence is constructed.

[0035] When constructing a conversation sequence for a multi-person voice call, if the conversation sequence is built based on all or most of the conversation content after the multi-person voice call is established, the final confirmation of the key points of the call will be inaccurate and inefficient, or even impossible.

[0036] Therefore, by considering the scenario of multi-person voice calls and the number of participants, the sequence window size of the conversation sequence to be constructed is determined. Then, based on the sequence window size of the conversation sequence to be constructed, the conversation sequence of multi-person voice calls is constructed, thereby avoiding the shortcomings of the above-mentioned conversation sequence without the limitation of the conversation content.

[0037] The more complex the call scenario and the more participants in the voice call, the larger the sequence window size should be within the threshold limit, meaning it cannot exceed the threshold size. In one embodiment, the range of different participants in different call scenarios and the corresponding sequence window size for each range can be pre-defined. This allows for a direct and quick determination of the sequence window size for the conversation sequence to be constructed based on the call scenario and the range of participants.

[0038] In one feasible implementation, step A11 may include steps A11a to A11b: Step A11a: Construct a prediction model for the sequence window size using historical data, where the historical data includes historical call scenarios, historical voice call participants, and the historical sequence window size of the historical conversation sequences corresponding to the historical call scenarios and historical voice call participants. Step A11b: Using a prediction model, determine the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of people in the voice call.

[0039] In this embodiment, another method is provided for determining the sequence window size of the session sequence to be constructed.

[0040] Because there is a clear correlation between the call scenario, the number of participants in the voice call, and the sequence window size of the corresponding conversation sequence—that is, the more complex the call scenario and the more participants, the larger the sequence window size will be within a threshold limit—a predictive model for the sequence window size can be constructed using historical data. This predictive model, based on the call scenario and the number of participants, can then determine the sequence window size for the conversation sequence to be constructed. Specifically, the call scenario and the number of participants can be converted into numerical features for model training and inference. Furthermore, the predictive model can be a tree model such as Random Forest, XGBoost, or LightGBM; a neural network model such as MLP or Embedding+fully connected layers; or a linear model such as Linear Regression or Lasso / Ridge.

[0041] In one feasible implementation, step A12 may include steps A12a to A12c: Step A12a: Obtain key historical calls within the historical time range and determine the risk level of the key historical calls; Step A12b: Adjust the sequence window size of the session sequence according to the risk level to obtain the latest sequence window size of the session sequence; Step A12c: Based on the latest sequence window size of the session sequence to be constructed, construct the session sequence for multi-person voice calls.

[0042] After determining the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of people in the voice call, the sequence window size of the conversation sequence can be adjusted according to the risk level corresponding to the key points of historical calls within the historical time range. Based on the latest sequence window size of the conversation sequence to be constructed, the conversation sequence of multi-person voice calls can be accurately constructed.

[0043] The risk level of historical call highlights can be determined by identifying specific keywords related to safety production risks within the call highlights. For example, the risk level can be determined by the quantity and frequency of specific keywords, or it can be inferred from these keywords. The higher the risk level of historical call highlights within a historical timeframe, the smaller the sequence window size of the conversation sequence, thereby improving the granularity of the conversation sequence. Furthermore, in high-risk emergency situations, the accuracy of the call highlights in multi-person voice calls can be improved by appropriately discarding the comprehensive coverage of the call highlights.

[0044] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 After step S20, the method for confirming the focus of the call further includes steps B11 to B14: Step B11: Divide each voice call participant into a task group and obtain the task scenario for each task group; Step B12: Construct the conversation sequence for each task group, wherein the conversation sequence for each task group includes the task scenario, the roles of each voice call participant in the task scenario, and the preset keywords of each voice call participant in the task scenario. Step B13: Identify the key points of each task group's calls through the conversation sequence of the task groups; Step B14: Determine the focus of the multi-person voice call based on the task scenario and call focus of each task group.

[0045] Each voice call participant can be divided into different task groups using clustering algorithms based on the task tags corresponding to the multiple tasks assigned to them. The task scenario of each task group can be inferred based on the role of each voice call participant in the task group. Alternatively, task groups under the target task scenario can be pre-defined, and each voice call participant can be divided into the corresponding task group based on their role.

[0046] A conversation sequence is constructed for each task group, consisting of a task scenario, the roles of each voice call participant in the task scenario, and preset keywords of each voice call participant in the task scenario. The conversation sequence of the task group is similar to the conversation sequence of the multi-person voice call in the first embodiment above. The conversation focus of each task group is identified through the conversation sequence of the task group, which is similar to the conversation focus of the multi-person voice call identified through the conversation sequence of the multi-person voice call in the first embodiment above. Therefore, it will not be described again here.

[0047] After identifying the call priorities of each task group, the call priorities of each task group carrying the task scenarios of each task group can be directly used as the call priorities of the multi-person voice call; alternatively, based on the priority of the task scenarios of each task group, the call priorities of some task groups with higher priority can be used as the call priorities of the multi-person voice call; alternatively, the call priorities of task groups with the same or similar task scenarios can be deduplicated and concatenated to obtain the call priorities of the multi-person voice call.

[0048] In addition to constructing the conversation sequence of each voice call participant in the call scenario in the first embodiment above to uniformly confirm the call focus of multi-person voice calls, it is also possible to divide each voice call participant into task groups and construct conversation sequences for each task group. The call focus of each task group can be identified through its conversation sequence, and then the call focus of the multi-person voice call can be determined based on the task scenario and call focus of each task group. This approach better aligns with actual safety production scenarios, more accurately identifies the call focus of each task group, and ultimately makes the call focus of the multi-person voice call determined based on the task scenario and call focus of each task group more accurate.

[0049] In one possible implementation, steps B14 may be followed by steps B14a to B14c: Step B14a: The key points of the multi-person voice call identified through the conversation sequence of the multi-person voice call are taken as the first key points of the call. Step B14b: The call focus of the multi-person voice call, determined according to the task scenario and call focus of each task group, shall be used as the second call focus. Step B14c: Determine the third call priority based on the first call priority and the second call priority, and use the third call priority as the target call priority for multi-person voice calls.

[0050] For the specific implementation of the call focus of a multi-person voice call obtained by recognizing the conversation sequence of the multi-person voice call, and the call focus of the multi-person voice call determined according to the task scenario and call focus of each task group, please refer to the above embodiment, which will not be repeated here.

[0051] After determining the first and second call priorities, a third call priority is determined based on the first and second call priorities. Specifically, the third call priority can be determined as the call priority with higher confidence based on the confidence level of the first call priority in a multi-person voice call scenario and the confidence level of the second call priority in a task grouping scenario. Alternatively, the first and second call priorities can be combined, and the same or similar call priorities can be merged to determine the third call priority as the combined call priority.

[0052] In this embodiment, the target call focus of the multi-person voice call is determined by combining the call focus of the multi-person voice call obtained by recognizing the conversation sequence of the multi-person voice call, and the call focus of the multi-person voice call determined by the task scenario and call focus of each task group, so that the determination of the final call focus is more accurate and has better robustness.

[0053] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method for confirming the focus of the call in this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0054] This application also provides a device for confirming call focus, please refer to... Figure 4 The call focus confirmation device includes: The conversion module 10 is used to acquire the call scenario of a multi-person voice call and the voice content of each participant in the voice call, and convert the voice content into text content. The determination module 20 is used to determine the preset keywords of each voice call participant in the call scenario based on the text content; The construction module 30 is used to construct a conversation sequence for a multi-person voice call. The conversation sequence includes a call scenario, the roles of each voice call participant in the call scenario, and preset keywords for each voice call participant in the call scenario. The recognition module 40 is used to identify the key points of a multi-person voice call through the conversation sequence of the multi-person voice call.

[0055] In one embodiment, the construction module 30 is further configured to: Obtain the number of participants in a multi-person voice call, and determine the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of participants. Based on the sequence window size of the session sequence to be constructed, a multi-person voice call session sequence is constructed.

[0056] In one embodiment, the construction module 30 is further configured to: A predictive model for sequence window size is constructed using historical data, which includes historical call scenarios, historical voice call participants, and historical sequence window sizes for historical conversation sequences corresponding to historical call scenarios and historical voice call participants. The sequence window size for the conversation sequence to be constructed is determined by using a predictive model, based on the call scenario and the number of people in the voice call.

[0057] In one embodiment, the construction module 30 is further configured to: Obtain key historical calls within a historical time frame and determine the risk level of those key historical calls; Adjust the sequence window size of the session sequence according to the risk level to obtain the latest sequence window size of the session sequence; Based on the latest sequence window size of the session sequence to be constructed, a multi-person voice call session sequence is constructed.

[0058] In one embodiment, the call focus confirmation device further includes a grouping module for: After the step of determining the preset keywords for each voice call participant in the call scenario based on the text content: Each participant in the voice call is assigned to a task group, and the task scenario for each task group is obtained. Construct a conversation sequence for each task group, wherein the conversation sequence for each task group includes the task scenario, the roles of each voice call participant in the task scenario, and the preset keywords of each voice call participant in the task scenario. By analyzing the conversation sequences of task groups, the key points of each task group's calls can be identified. Based on the task scenarios and call priorities of each task group, determine the key points of multi-person voice calls.

[0059] In one embodiment, the grouping module is further configured to: After the step of determining the focus of a multi-person voice call based on the task scenario and call focus of each task group: The key points of a multi-person voice call, identified through the conversation sequence of the multi-person voice call, are taken as the first key point of the call. The call priorities for multi-person voice calls, determined based on the task scenarios and call priorities of each task group, will be used as the second call priorities. The third call priority is determined based on the first and second call priorities, and this third call priority is used as the target call priority for multi-person voice calls.

[0060] The call focus confirmation device provided in this application, employing the call focus confirmation method in the above embodiments, can solve the technical problems of incomplete, inefficient, and inaccurate capture and confirmation of call focus in multi-person voice calls. Compared with the prior art, the beneficial effects of the call focus confirmation device provided in this application are the same as those of the call focus confirmation method provided in the above embodiments, and other technical features in the call focus confirmation device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0061] This application provides a call focus confirmation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the call focus confirmation method in the first embodiment described above.

[0062] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a call focus confirmation device suitable for implementing the embodiments of this application. The call focus confirmation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The device shown for confirming call focus is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0063] like Figure 5As shown, the call focus confirmation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the call focus confirmation device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the call confirmation device to communicate wirelessly or wiredly with other devices to exchange data. Although call confirmation devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0064] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0065] The call focus confirmation device provided in this application, employing the call focus confirmation method in the above embodiments, can solve the technical problems of incomplete, inefficient, and inaccurate capture and confirmation of call focus in multi-person voice calls. Compared with the prior art, the beneficial effects of the call focus confirmation device provided in this application are the same as those of the call focus confirmation method provided in the above embodiments, and other technical features of this call focus confirmation device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0066] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0067] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0068] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the call focus confirmation method in the above embodiments.

[0069] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0070] The aforementioned computer-readable storage medium may be included in the call focus confirmation device; or it may exist independently and not be assembled into the call focus confirmation device.

[0071] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a call focus confirmation device, cause the call focus confirmation device to: acquire the call scenario of a multi-person voice call and the voice content of each voice call participant, and convert the voice content into text content; determine the preset keywords of each voice call participant in the call scenario based on the text content; construct a conversation sequence of the multi-person voice call, wherein the conversation sequence includes the call scenario, the roles of each voice call participant in the call scenario, and the preset keywords of each voice call participant in the call scenario; and identify the call focus of the multi-person voice call through the conversation sequence of the multi-person voice call.

[0072] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0074] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0075] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for confirming call priorities. This solves the technical problems of incomplete, inefficient, and inaccurate capture and confirmation of call priorities in multi-person voice calls. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the call priority confirmation method provided in the above embodiments, and will not be repeated here.

[0076] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the call focus confirmation method described above.

[0077] The computer program product provided in this application can solve the technical problems of incomplete, inefficient, and inaccurate capture and confirmation of key points in multi-person voice calls. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the call key point confirmation method provided in the above embodiments, and will not be repeated here.

[0078] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for confirming the focus of a call, characterized in that, The methods for confirming the focus of the call include: The system acquires the call scenario of a multi-person voice call and the voice content of each participant, and converts the voice content into text content. Based on the text content, determine the preset keywords for each voice call participant in the call scenario; Construct a conversation sequence for a multi-person voice call, wherein the conversation sequence includes the call scenario, the roles of each voice call participant in the call scenario, and preset keywords for each voice call participant in the call scenario; By analyzing the conversation sequence of a multi-person voice call, the key points of the call can be identified.

2. The method for confirming the focus of a call as described in claim 1, characterized in that, The steps for constructing the conversation sequence for a multi-person voice call include: Obtain the number of participants in a multi-person voice call, and determine the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of participants. Based on the sequence window size of the session sequence to be constructed, a multi-person voice call session sequence is constructed.

3. The method for confirming the focus of a call as described in claim 2, characterized in that, The step of determining the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of people in the voice call includes: A prediction model for the sequence window size is constructed using historical data, wherein the historical data includes historical call scenarios, historical voice call participants, and the historical sequence window size of the historical conversation sequences corresponding to the historical call scenarios and the historical voice call participants. The prediction model determines the sequence window size of the conversation sequence to be constructed based on the call scenario and the number of people in the voice call.

4. The method for confirming the focus of a call as described in claim 2, characterized in that, The step of constructing a multi-person voice call conversation sequence based on the sequence window size of the conversation sequence to be constructed includes: Obtain key historical calls within a historical time range and determine the risk level of those key historical calls; Adjust the sequence window size of the session sequence according to the risk level to obtain the latest sequence window size of the session sequence; Based on the latest sequence window size of the session sequence to be constructed, a multi-person voice call session sequence is constructed.

5. The method for confirming the focus of a call as described in claim 1, characterized in that, After the step of determining the preset keywords of each voice call participant in the call scenario based on the text content, the method for confirming the call focus further includes: Each participant in the voice call is assigned to a task group, and the task scenario for each task group is obtained. Construct a conversation sequence for each task group, wherein the conversation sequence for each task group includes the task scenario, the roles of each voice call participant in the task scenario, and the preset keywords of each voice call participant in the task scenario. By analyzing the conversation sequences of task groups, the key points of each task group's calls can be identified. Based on the task scenarios and call priorities of each task group, determine the key points of multi-person voice calls.

6. The method for confirming the focus of a call as described in claim 5, characterized in that, Following the step of determining the call focus of a multi-person voice call based on the task scenario and call focus of each task group, the following further steps are included: The key points of a multi-person voice call, identified through the conversation sequence of the multi-person voice call, are taken as the first key point of the call. The call priorities for multi-person voice calls, determined based on the task scenarios and call priorities of each task group, will be used as the second call priorities. A third call focus is determined based on the first and second call focus, and this third call focus is used as the target call focus for multi-person voice calls.

7. A device for confirming the focus of a call, characterized in that, The device for confirming the call focus includes: The conversion module is used to acquire the call scenario of a multi-person voice call and the voice content of each participant in the voice call, and convert the voice content into text content. The determination module is used to determine the preset keywords of each voice call participant in the call scenario based on the text content. A construction module is used to construct a conversation sequence for a multi-person voice call, wherein the conversation sequence includes the call scenario, the roles of each voice call participant in the call scenario, and the preset keywords of each voice call participant in the call scenario; The recognition module is used to identify the key points of a multi-person voice call by analyzing the conversation sequence of the multi-person voice call.

8. A device for confirming call focus, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the call focus confirmation method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the call focus confirmation method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the call focus confirmation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Structured content processing method and device for multi-person conference scene, equipment and medium

    CN109783642A

  • Voice call management method, device, equipment, storage medium and product

    CN119788121A

  • Poisonous bait station specification determination method, system and equipment based on poison bait station positioning

    CN119886844A

  • Input method and apparatus and electronic device

    IN201747012105A

  • Task processing method, task processing model training method, and conference speech separation method

    WO2025092405A1