Scheduling method and device and storage medium

By using the voice to text model and named entity recognition model in the first aid dispatching system, the key information in the first aid phone is automatically recognized and input, which solves the problem of cumbersome information manually entering the dispatcher, and improves the accuracy and efficiency of scheduling.

CN119993405APending Publication Date: 2025-05-13CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311498726.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the first aid dispatch system, the dispatcher needs to manually record the incident and complaint information, which leads to cumbersome and complicated work, which may affect the accuracy of the dispatch.

Method used

The received voice information is converted into text information using the first model, and the first type of entity information (including at least address information and symptom information) in the text information is determined through the second model, and the information is input to the dispatching station.

Benefits of technology

Simplify the workflow of the dispatcher, improve the accuracy and efficiency of the dispatch, and reduce errors and delays caused by manual entry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993405A_ABST
    Figure CN119993405A_ABST
Patent Text Reader

Abstract

The invention discloses a scheduling method and device and a storage medium. The scheduling method comprises the following steps: converting received voice information into text information by adopting a first model; adopting a second model to determine first-class entity information and second-class entity information in the text information; wherein the first type of entity information at least comprises first address information; and at least inputting the first type of entity information to a dispatching desk. Thus, the first type of entity information and the second type of entity information are determined through the first model and the second model, and at least the first type of entity information including the first address information is input to the dispatching desk, so that the work of a dispatcher is simplified, the work convenience degree of the dispatcher is increased, dispatching can be performed according to the first address information, and the dispatching efficiency is improved. Therefore, the scheduling accuracy of the scheduling station is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vertical industry technology, and in particular to a scheduling method, device and storage medium. Background Art

[0002] In the current emergency dispatch system, the answering of emergency calls and the dispatching of ambulances are mainly completed by emergency dispatchers, so emergency dispatchers serve as a bridge of communication between patients and ambulances. Every minute and every second of the 120 emergency dispatcher is racing against time, and life is precious and cannot be repeated.

[0003] However, there are still some problems in 120 emergency dispatch. For example, after receiving an emergency call, the dispatcher needs to listen to the caller's call for help during the call, and also needs to register and record the event and the main complaint and save the stub for subsequent processing and use. The registration and preservation of events and main complaints are mainly input manually by the dispatcher, which makes the dispatcher's work cumbersome and complicated, and may affect the accuracy of the dispatch desk's emergency dispatch. Summary of the invention

[0004] In order to solve the existing technical problems, the embodiments of the present invention provide a scheduling method, device and storage medium.

[0005] The technical solution of the present invention is achieved in this way:

[0006] An embodiment of the present invention provides a scheduling method, the method comprising:

[0007] Converting the received voice information into text information using the first model;

[0008] Using the second model, determining the first type of entity information and the second type of entity information in the text information; wherein the first type of entity information at least includes: first address information;

[0009] At least the first type of entity information is input into the dispatching console.

[0010] In the above solution, the first type of entity information at least includes: symptom information.

[0011] In the above scheme, the method further includes:

[0012] Performing data enhancement on the speech training data, and using the data enhanced speech training data to train the first model, wherein the data enhancement includes at least one of the following: volume disturbance, speech rate disturbance, and noise disturbance.

[0013] In the above scheme, the method further includes:

[0014] Determining a second loss function corresponding to the first model according to speech training data associated with the first type of entity information;

[0015] Determining a third loss function corresponding to the first model according to speech training data associated with the second type of entity information;

[0016] A first loss function corresponding to the first model is determined according to the second loss function and the third loss function.

[0017] In the above solution, determining the first loss function corresponding to the first model according to the second loss function and the third loss function includes:

[0018] The first loss function corresponding to the first model is expressed by the following formula:

[0019] L all =λ 1 L(S address )+λ 2 L(S symptom )+λ 3 L(S other )

[0020] Among them, L(S address ) represents the second loss function corresponding to the first address information, L(S symptom ) represents the second loss function corresponding to symptom information, L(S other ) represents the third loss function corresponding to the second type of entity information, λ 1 represents the first weight of the second loss function corresponding to the first address information, λ 2 represents the second weight of the second loss function corresponding to the symptom information, λ 3 Represents a third weight of a third loss function corresponding to the second type of entity information; the first weight is greater than the third weight, and the second weight is greater than the third weight.

[0021] In the above solution, the adopting the second model to determine the first type of entity information and the second type of entity information in the text information includes:

[0022] Using the second model to perform a binary classification task on the text information to determine the first type of entity information in the text information;

[0023] The second model is used to perform N classification tasks on the text information to determine the second type of entity information in the text information, where N is an integer greater than or equal to 2.

[0024] In the above scheme, the method further includes:

[0025] Determine at least one second address information of a sending end of the voice information according to a preset positioning method;

[0026] determining a target address according to the first address information and the at least one second address information;

[0027] At least the first type of entity information including the target address is input into the dispatching console.

[0028] The embodiment of the present invention provides a scheduling device, which includes: a conversion module, a first determination module, and a processing module; wherein:

[0029] The conversion module is used to convert the received voice information into text information using the first model;

[0030] The first determination module is used to determine the first type of entity information and the second type of entity information in the text information by using the second model; wherein the first type of entity information at least includes: first address information;

[0031] The processing module is used to input at least the first type of entity information into the dispatching console.

[0032] In the above solution, the first type of entity information at least includes: symptom information.

[0033] In the above solution, the device further includes: a first training module; wherein,

[0034] The first training module is used to perform data enhancement on speech training data, and use the data enhanced speech training data to train the first model, wherein the data enhancement includes at least one of the following: volume disturbance, speech rate disturbance and noise disturbance.

[0035] In the above solution, the device further includes: a second training module; wherein,

[0036] The second training module is used to determine a second loss function corresponding to the first model based on the speech training data associated with the first type of entity information;

[0037] Determining a third loss function corresponding to the first model according to speech training data associated with the second type of entity information;

[0038] A first loss function corresponding to the first model is determined according to the second loss function and the third loss function.

[0039] In the above solution, the second training module is specifically used to express the first loss function corresponding to the first model using the following formula:

[0040] L all =λ1 L(S address )+λ 2 L(S symptom )+λ 3 L(S other )

[0041] Among them, L(S address ) represents the second loss function corresponding to the first address information, L(S symptom ) represents the second loss function corresponding to symptom information, L(S other ) represents the third loss function corresponding to the second type of entity information, λ 1 represents the first weight of the second loss function corresponding to the first address information, λ 2 represents the second weight of the second loss function corresponding to the symptom information, λ 3 Represents a third weight of a third loss function corresponding to the second type of entity information; the first weight is greater than the third weight, and the second weight is greater than the third weight.

[0042] In the above solution, the first determination module is specifically used to perform a binary classification task on the text information using the second model to determine the first type of entity information in the text information;

[0043] The second model is used to perform N classification tasks on the text information to determine the second type of entity information in the text information, where N is an integer greater than or equal to 2.

[0044] In the above solution, the device further includes: a second determination module; wherein,

[0045] The second determination module is used to determine at least one second address information of the sender of the voice information according to a preset positioning method;

[0046] determining a target address according to the first address information and the at least one second address information;

[0047] At least the first type of entity information including the target address is input into the dispatching console.

[0048] In a third aspect, an embodiment of the present invention further provides a scheduling device, comprising a memory, a processor, and an executable program stored in the memory and capable of being run by the processor, wherein the processor executes any step of the scheduling method when running the executable program.

[0049] In a fourth aspect, an embodiment of the present invention further provides a storage medium having an executable program stored thereon, wherein the executable program, when executed by a processor, implements the steps of any one of the scheduling methods.

[0050] The dispatching method, device, and computer-readable storage medium provided in the embodiments of the present invention use a first model to convert received voice information into text information; use a second model to determine the first type of entity information and the second type of entity information in the text information; wherein the first type of entity information includes at least: first address information; at least the first type of entity information is input into the dispatching desk; in this way, the first type of entity information and the second type of entity information are determined by the first model and the second model, and at least the first type of entity information including the first address information is input into the dispatching desk, thereby increasing the convenience of the dispatcher's work, and can also directly perform dispatch according to the first address information, thereby improving the accuracy of dispatching by the dispatching desk. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of a scheduling method provided by an embodiment of the present invention;

[0052] Figure 2 A schematic diagram of a second model structure of a scheduling method provided by an embodiment of the present invention;

[0053] Figure 3 A schematic diagram of a flow chart of another scheduling method provided by an embodiment of the present invention;

[0054] Figure 4 A flowchart of another scheduling method provided by an embodiment of the present invention;

[0055] Figure 5 A schematic diagram of a target address determination process of a scheduling method provided by an embodiment of the present invention;

[0056] Figure 6 A schematic diagram of the structure of a scheduling device provided by an embodiment of the present invention;

[0057] Figure 7 A schematic diagram of the structure of another scheduling device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] At present, in the field of input systems of existing emergency dispatch desks, there are the following problems:

[0059] First, the event record needs to be manually recorded by the dispatcher, which restricts the dispatcher's hands and energy and prevents him from directly rescuing the patient on the scene, thus wasting the patient's precious rescue time.

[0060] Second, when the dispatcher is listening and recording, because the patient's speaking speed is faster than the dispatcher's typing speed, there will be repeated inquiries for key information, wasting precious rescue time;

[0061] Third, in the actual dispatch process, the dispatcher cannot obtain the location and symptom information in a timely and accurate manner due to the lack of clarity expressed by the caller, resulting in a delay in the rescue time;

[0062] Fourth, manual record keeping may result in the loss of key information, which is not conducive to the reuse and retention of information and may easily lead to certain medical disputes.

[0063] To solve at least one of the above problems, various embodiments of the present invention provide a scheduling method, including: using a first model to convert received voice information into text information; using a second model to determine first-category entity information and second-category entity information in the text information; and inputting the first-category entity information and the second-category entity information into a scheduling console.

[0064] The present invention will be further described in detail below in conjunction with the embodiments.

[0065] Figure 1 A flowchart of a scheduling method provided by an embodiment of the present invention; Figure 1 As shown, the method includes:

[0066] Step 101: using a first model to convert received voice information into text information;

[0067] Step 102: using a second model, determining first-category entity information and second-category entity information in the text information; wherein the first-category entity information at least includes: first address information;

[0068] Step 103: input at least the first type of entity information into the dispatching console.

[0069] In a possible implementation, the method is applied to a scheduling system; the scheduling system is used to execute the method; the scheduling system may include: a first model, a second model, and a scheduling console.

[0070] In a possible implementation, the first model and the second model may be machine learning models. The first model may be a speech-to-text (ASR, Automatic Speech Recognition) model, the input of the model may be speech information, and the output may be text information; the second model may be a named entity recognition (NER, Named Entity Recognition) model (semantic recognition classification), which can automatically identify the first type of entity information and the second type of entity information in the text information.

[0071] Here, entity (or semantics) can refer to specific, actually existing objects or things; entity information can refer to key information related to business scenarios; named entity recognition can refer to identifying entities with specific meanings in text information, which can include but are not limited to: names of people, places, organizations, proper nouns, etc.

[0072] For example, when the business scenario is railway transport scheduling, the entity information may include but is not limited to: name, flight number, departure time, and stops along the way; when the business scenario is emergency center scheduling, the entity information may include but is not limited to: name, address, symptoms, and age.

[0073] In a possible implementation, the first type of entity information may be necessary entity information, and the second type of entity information may be non-necessary entity information; the necessary entity information may refer to entity information with a higher priority.

[0074] Here, with respect to the priority input by the dispatching station, the priority of the first type of entity information may be higher than the priority of the second type of entity information.

[0075] In one possible implementation, the dispatch desk may refer to a control center or a command and management center, through which data can be monitored and instructions can be issued; the dispatch desk may include but is not limited to: a power dispatch desk, a nuclear industry dispatch desk, an aerospace dispatch desk, an emergency center dispatch desk, etc.; the dispatch desk can be commanded by a dispatcher.

[0076] Specifically, voice information is received, and the first model is used to convert the voice information into text information; the second model is used to identify entity information in the text information, and the first type of entity information and the second type of entity information are determined; the first type of entity information and the second type of entity information are input into the dispatching console for dispatching.

[0077] In one possible implementation, at least the first type of entity information can be input into the dispatching console, and business scheduling is first performed based on the first type of entity information. Then, the second type of entity information is input into the dispatching console, and the dispatching console combines the received first type of entity information with the second type of entity information for further business scheduling.

[0078] Here, only the first type of entity information is input into the dispatching console, and business dispatching can also be performed.

[0079] In actual applications, if the dispatch desk is an airport dispatch desk, it will dispatch the airport runway according to the received flight departure time, and promptly clear out irrelevant personnel in preparation for takeoff; if the dispatch desk is an emergency center dispatch desk, it will dispatch ambulances and medical personnel according to the received emergency information such as patient symptoms and onset address, and promptly go to the location for rescue; when dispatching an ambulance, the emergency center dispatch desk can give priority to informing the medical personnel of critical emergency information, and the rest of the non-critical emergency information (such as the patient's age and gender) can be informed after the ambulance has departed, which can improve the dispatch efficiency.

[0080] Through the above method, the first model and the second model are used to convert voice information into text information and identify entity information in the text information, and the entity information is divided into first-category entity information and second-category entity information, and at least the first-category entity information is input into the dispatch desk, so that the dispatcher no longer needs to manually input the entity information, which simplifies the dispatcher's workflow and increases the dispatcher's work convenience; and because the second model can identify the first-category entity information including at least the first address information, it can be convenient for the dispatcher to directly dispatch medical resources according to the first address information, greatly increasing the efficiency of the dispatch desk.

[0081] In some embodiments, the first category of entity information also includes at least: symptom information.

[0082] Specifically, the second model is used to preferentially identify the first type of entity information in the text information, including the first address information and the symptom information, and then determine the second type of entity information in the text information.

[0083] In one possible implementation, the first address information may include at least one of the following: road name, surrounding landmark buildings, floor, etc.; the symptom information may include at least one of the following: onset time, patient's current condition, symptom severity, etc.

[0084] Through the above method, the dispatch desk can accurately identify the first address information and symptom information in the corresponding text information of the caller, and directly dispatch doctors and ambulances based on the first address information and symptom information, avoiding wasting rescue time due to unclear expression of the caller, and at the same time improving the professionalism and accuracy of medical dispatch.

[0085] In some embodiments, the method further comprises:

[0086] Performing data enhancement on the speech training data, and using the data enhanced speech training data to train the first model, wherein the data enhancement includes at least one of the following: volume disturbance, speech rate disturbance, and noise disturbance.

[0087] Specifically, before training the first model according to the speech training data, data enhancement is performed on the speech training data to increase the robustness of the first model, and the data enhancement includes: volume disturbance, speech rate disturbance and noise disturbance.

[0088] In one possible implementation, the volume disturbance may be to randomly select a volume coefficient within a preset volume value range to adjust the volume of the speech training data to enhance the robustness of the model to the volume size; the preset volume value range may be [0.7, 1.4], which may be specifically set in actual applications and is not specifically limited here.

[0089] Here, the volume of the speech training data is adjusted by randomly selecting a volume coefficient within a preset volume value interval, and the volume coefficient of the model is determined when the robustness to the volume is the highest.

[0090] In one possible implementation, the speech rate disturbance may be to randomly select a speech rate coefficient within a preset speed value range to adjust the speech speed of the speech training data to enhance the robustness of the model to the speed of speech; the preset speed value range may be [0.8, 1.2], which may be specifically set in actual applications and is not specifically limited here.

[0091] Here, the speech speed of the speech training data is adjusted by randomly selecting a speech speed coefficient within a preset speed value range, and the speech speed coefficient of the model is determined when the model has the highest robustness to the speech speed.

[0092] In one possible implementation, noise interference can be to randomly select a noise size within a preset noise value range, and add noise to all speech training data to enhance the generalization of the model to noisy speech data, so that the model can cover more scenarios; the preset noise value range can be [10dB, 500dB], which can be specifically set in actual applications and is not specifically limited here.

[0093] By using the above method, data enhancement is performed on the voice training data when training the first model, which can reduce the interference of environmental noise in the process of converting the voice information of the caller into text information, thereby improving the recognition accuracy and robustness of the dispatch station.

[0094] In some embodiments, the method further comprises:

[0095] Determining a second loss function corresponding to the first model according to speech training data associated with the first type of entity information;

[0096] Determining a third loss function corresponding to the first model according to speech training data associated with the second type of entity information;

[0097] A first loss function corresponding to the first model is determined according to the second loss function and the third loss function.

[0098] Specifically, part of the speech training data that is associated with the first address information and symptom information is separately marked to determine the second loss function corresponding to the first model; the third loss function corresponding to the first model is determined based on the speech training data corresponding to the second type of entity information; and the first loss function corresponding to the first model is determined based on the second loss function and the third loss function.

[0099] In a possible implementation, the speech training data may be trimmed, and the speech training data associated with the first address information and the symptom information in the speech training data may be trimmed, and used as separate speech training data to train the model.

[0100] Through the above method, the first model can be made more sensitive to the first address information and symptom information in the process of converting the voice information of the caller into text information, thereby enhancing the recognition accuracy and efficiency of the first model for the first address information and symptom information.

[0101] In some embodiments, determining the first loss function corresponding to the first model according to the second loss function and the third loss function includes:

[0102] The first loss function corresponding to the first model is expressed by the following formula:

[0103] L all =λ 1 L(S address )+λ 2 L(S symptom )+λ 3 L(S other )

[0104] Among them, L(S address ) represents the second loss function corresponding to the first address information, L(S symptom ) represents the second loss function corresponding to symptom information, L(S other ) represents the third loss function corresponding to the second type of entity information, λ 1 represents the first weight of the second loss function corresponding to the first address information, λ 2 represents the second weight of the second loss function corresponding to the symptom information, λ 3 Represents a third weight of a third loss function corresponding to the second type of entity information; the first weight is greater than the third weight, and the second weight is greater than the third weight.

[0105] In one possible implementation, the relationship between the first weight, the second weight, and the third weight can be set so that the first weight is greater than the third weight, and the second weight is also greater than the third weight. In this way, the first model can be more sensitive to the recognition of the first location information and symptom information.

[0106] For example, 1 , 2 With λ 3 can be a non-negative decimal, and λ 1 , 2 With λ 3 can sum to 1; for example, λ 1 The value of can be 0.4, λ 2 The value of can be 0.4, λ 3 The value of can be 0.2; the values ​​of the first weight, the second weight and the third weight can be set according to actual needs and are not limited here.

[0107] In a possible implementation, the second type of entity information may include at least one of the following: the relationship between the person calling for help and the patient, whether the person calling for help is with the patient, the patient's age, the patient's name, whether the patient is breathing, and whether the patient is conscious.

[0108] Specifically, the speech training data associated with the first address information and the speech training data associated with the symptom information in the speech training data are separately marked, and the second loss function corresponding to the first address information, the second loss function corresponding to the symptom information, and the third loss function corresponding to the second type of entity information are respectively determined, and weight values ​​are respectively assigned to the second loss function and the third loss function, and the weight value corresponding to the minimum loss value between the output data of the loss function and the speech training data is determined, and the loss function at this time is used as the first loss function; according to the first loss function, the first model is determined.

[0109] In one possible implementation, the second loss function and the third loss function may include but are not limited to: a connectionist temporal classification (CTC) loss function, a cross entropy loss function, and a recurrent neural network transducer (RNN-T) loss function.

[0110] In a possible implementation, the second loss function and the third loss function may be CTC loss functions, which may be expressed by the following formula:

[0111]

[0112] Among them, S can represent the speech training data set, which is a subset of the overall distribution; x can be the prediction result of the training model; z can be the true label corresponding to x; p(z|x) can represent the probability of restoring x to z, that is, the conditional probability of z relative to x; Can be a negative log-likelihood function.

[0113] Here, the negative log-likelihood function can be a loss function used to solve classification problems. It is a natural logarithm form of the likelihood function, which can be used to measure the similarity between two probability distributions. The negative sign is taken to make the maximum likelihood value correspond to the minimum loss.

[0114] In one possible implementation, CTC can be an algorithm for speech recognition and text recognition. For a given input sequence x, CTC can calculate the probability distribution of all possible output sequences z. Through the probability distribution, the output corresponding to the maximum probability or the probability of a predetermined output can be predicted. The predetermined output can be set according to specific needs.

[0115] For example, the minimum value of the loss function L(S) can be determined by minimizing the negative log-likelihood function By iterating the loss function, the predicted result x can be continuously approached to the true label z, so that the predicted result of the training model is closer to the true result.

[0116] In one possible implementation, the speech training data can be divided into a training set, a validation set, and a test set according to a predetermined ratio; the model can be trained using the training set until the model converges; the converged model can be tested using the validation set and the test set, and the optimal model among the converged models can be determined as the first model.

[0117] Exemplarily, the predetermined ratio may be 6:2:2, which may be set according to actual needs.

[0118] Through the above method, the first model is determined according to the first loss function, and the voice information is converted into text information according to the first model, which can improve the conversion accuracy of the first model for the first address information and symptom information, and avoid delays in rescue time due to inaccurate information acquisition.

[0119] In some embodiments, the using the second model to determine the first type of entity information and the second type of entity information in the text information includes:

[0120] Using the second model to perform a binary classification task on the text information to determine the first type of entity information in the text information;

[0121] The second model is used to perform N classification tasks on the text information to determine the second type of entity information in the text information, where N is an integer greater than or equal to 2.

[0122] Specifically, a second model is determined according to the speech training data, a binary classification task is used in the second model to determine the first type of entity information in the text information, and an N-classification task is used to determine the second type of entity information in the text information.

[0123] Here, the binary classification task may be a classification method for determining whether information is or is not a specific content; the N-classification task may be a classification method for determining whether information is a specific type of N specific contents.

[0124] Exemplarily, the training model may be text information of a voice call record of an emergency center dispatch desk, and entity information may be marked in the text information.

[0125] For example, the text information may be "Doctor, please come quickly, our elderly has chest pain and sweating after going to the toilet, please come and save him! The address is XX address, X building, X unit, XXX", and the entity information marked in the text information may be "symptom information: chest pain, sweating, first address information: XX address, X building, X unit, XXX, relationship with the patient: our elderly (i.e., relative), gender: None and age: None".

[0126] Here, when specific entity information cannot be obtained from the text information, "None" can be filled in the corresponding position of the dispatch console.

[0127] In one possible implementation, when determining the second model, the common named entity recognition model in the prior art can be improved, and the first position information and symptom information in the speech training data can be trained separately using a parallel network.

[0128] In one possible implementation, the named entity recognition model can be based on sequence labeling. The segmented text information is labeled using common labeling rules such as the three-digit sequence labeling method (BIO, B-begin, I-inside, O-outside) and the four-digit sequence labeling method (BIOES, B-begin, M-middle, E-end, S-single). The model is constructed to predict each label in the text information, thereby performing entity recognition.

[0129] For example, in the BIO annotation rule, B can represent the beginning of an entity, I can represent the middle or end of an entity, and O can represent an entity that does not belong to the above B and O types; for example, in "My name is Zhang San", the annotation of "I" is "O", the annotation of "Jia" is "O", the annotation of "Zhang" is "B", and the annotation of "San" is "I".

[0130] In one possible implementation, the second model can use a natural language processing model (BERT, Bidirectional Encoder Representations from Transformers) and a bidirectional long short-term memory network (Bi-LSTM, Bi-directional Long Short-Term Memory) as the text feature editor at the bottom of the model structure, and perform entity label prediction through a fully connected conditional random field (FC-CRF, Fully Connected Conditional Random Field).

[0131] Exemplarily, the text information is input into the second model, and the text information can be converted into a vector through the BERT layer. The vector can build context information through the Bi-LSTM layer, and then enter the branch layer composed of FC-CRF for classification. A binary classification task can be used to determine the first address information and symptom information in the text information, and an N-classification task can be used to determine the second type of entity information in the text information.

[0132] For example, the structure of the second model can be as follows Figure 2 As shown, Figure 2 A schematic diagram of a second model structure of a scheduling method provided by an application embodiment of the present invention; Figure 2 As shown, the second model includes:

[0133] Input layer: The input layer contains...x n 、x n+1 、x n+2 、x n+3 ...may refer to input text information.

[0134] Encoding layer: BERT at the encoding layer can convert the input text information into text vector information.

[0135] Bi-LSTM layer: can build context information of text vector information.

[0136] Branching layer: The first address information and symptom information can be branched separately in the branching layer, and other second-category entity information can be branched. FC-CRF is used to determine whether the text vector information is the first address information, the symptom information, or other second-category entity information.

[0137] The first address information and symptom information in the text vector information are determined by adopting a binary classification task. The corresponding output in the first address information branch is "is the first address information" or "is not the first address information"; the corresponding output in the symptom information branch is "is the symptom information" or "is not the symptom information".

[0138] Here, in the branch layer, the type corresponding to the text vector information is determined by identifying the respective sequence annotations in each text vector information.

[0139] Exemplarily, in the position branch layer, the sequence annotation used to represent the position entity in the text vector information is judged to determine whether the text vector information is the first position information.

[0140] Furthermore, the way of adopting a parallel network can be to input the text vector information into each layer of the branch layer through at least one input path; each layer in the branch layer simultaneously judges the entity type of the text vector information, and outputs the type corresponding to the text vector information through at least one output path.

[0141] Output layer: The...y in the output layer can be used to... n , y n+1 , y n+2 , y n+3 ... Output the first type of entity information and the second type of entity information in the text information.

[0142] Here, the data format of the output layer can include but is not limited to an array, and the data of the output layer can be the sequence annotations corresponding to each character in the text information.

[0143] In a possible implementation, the BERT layer can convert the text information into vector information through the encoder and decoder of the Transformers package.

[0144] Exemplarily, the format of the input layer can be in the form of an array, and the elements in the array can be in the form of split characters; for example, the text information "My name is Zhang San" can be split, with "我 (wǒ)" as x 1 , "叫 (jiào)" as x 2 , "张 (zhāng)" as x 3 , "三 (sān)" as x 4 .

[0145] In a possible implementation, the input layer can also include the sequence annotation corresponding to each character.

[0146] Exemplarily, taking the BIO annotation rule as an example, in the input layer, "我 (wǒ)" and the sequence O are used as x 1 , "叫 (jiào)" and the sequence O are used as x 2 , "张 (zhāng)" and the sequence B are used as x 3 , "三 (sān)" and the sequence I are used as x 4 .

[0147] In one possible implementation, a Long Short-Term Memory (LSTM) layer can integrate contextual information to avoid the problem that when the vector sequence is long enough, it is difficult to transfer information from an earlier time step to a later time step.

[0148] Furthermore, multi-layer LSTM layers may refer to stacking LSTM layers, which may express the features of the vector more abstractly at a high level, increase the accuracy of model recognition, and reduce the model training time.

[0149] In a possible implementation, the Bi-LSTM may be composed of a forward LSTM and a backward LSTM; both the forward LSTM and the backward LSTM may be used to construct context information.

[0150] Here, the Bi-LSTM layer can better capture bidirectional semantic dependencies; for example, the word “not so” in “this restaurant is so dirty” is a modification of “dirty”, and using Bi-LSTM can better connect the contextual information.

[0151] In a possible implementation, by combining Bi-LSTM with CRF, it is possible to ensure that sufficient whole sentence features can be extracted while using an effective sequence labeling method for labeling.

[0152] For example, after the text vector information is input and passes through Bi-LSTM, the forward and backward hidden state results will be combined, which can effectively save the front and back information of the entire sentence text vector information, extract the feature information in the sentence, generate the output of Bi-LSTM, and use the output of Bi-LSTM as the input of CRF. CRF can use the context information to identify the corresponding sequence annotations in the text vector information, thereby improving the recognition accuracy.

[0153] In a possible implementation, N in the N classification tasks may be the same as the number of second-category entity information.

[0154] Exemplarily, when the second category of entity information includes the relationship with the patient, gender and age, the N classification task can be a 3-classification task, and the corresponding output in the three-classification task branch is one of the following: "is the relationship with the patient", "is the gender", "is the age".

[0155] In one possible implementation, the speech training data can be divided into a training set, a validation set, and a test set according to a predetermined ratio; the model can be trained with the training set until the model converges; the converged model can be tested with the validation set and the test set, and the optimal model among the converged models is determined as the second model.

[0156] Exemplarily, the predetermined ratio may be 6:2:2, which may be set according to actual needs.

[0157] Through the above method, the recognition effect of the first address information and symptom information can be enhanced, and excessive network computing costs can be avoided; through the second model, the entity information corresponding to the first address information and symptom information can be more efficiently identified, thereby meeting the needs of emergency dispatch.

[0158] In some embodiments, the method further comprises:

[0159] Determine at least one second address information of a sending end of the voice information according to a preset positioning method;

[0160] determining a target address according to the first address information and the at least one second address information;

[0161] At least the first type of entity information including the target address is input into the dispatching console.

[0162] In a possible implementation, the sender of the voice information may be the person calling for help.

[0163] Specifically, after determining the first address information, the first address information is input into the integrated positioning system to determine at least one second address information, and the target address is determined according to the first address information, the second address information and the corresponding weights.

[0164] Exemplarily, the preset positioning methods in the integrated positioning system may include, but are not limited to: entity recognition location, Global Positioning System (GPS) satellite positioning, Beidou satellite positioning, base station positioning and WIFI positioning.

[0165] In a possible implementation, the corresponding second address information is determined respectively by various preset positioning methods, the first address information and each second address information are voted, and the address information with the highest number of votes is used as the actual location.

[0166] Here, voting may refer to assigning different weight values ​​to each preset positioning method; for example, if there are five preset positioning methods (which may be entity recognition location, GPS satellite positioning, Beidou satellite positioning, base station positioning, and WIFI positioning), the weight values ​​may include λ 1 , 2 , 3 , 4 and λ 5 .

[0167] Exemplarily, the weight value may be determined according to the reliability of different preset positioning methods; for example, if the reliability of GPS satellite positioning is greater than that of WIFI positioning, the weight value λ of GPS satellite positioning is 2 Greater than the weight value λ of WIFI positioning 5 .

[0168] In one possible manner, the address information determined according to different preset positioning methods may be the same. The weight values ​​of the preset positioning methods with the same address information are added together, and finally the address information corresponding to the preset positioning method with the highest weight value is determined as the target address, and at least the first type of entity information containing the target address is input into the dispatching console; if the weight values ​​are the same, they can be determined through manual processing by the dispatcher.

[0169] Through the above method, the target address of the caller can be accurately obtained even if the caller does not express it clearly and input into the dispatch desk as the first type of entity information, thereby improving the accuracy and professionalism of the dispatch desk and avoiding affecting the dispatch speed of the ambulance due to inaccurate address.

[0170] In a possible implementation, after the first-category entity information and the second-category entity information are input into the dispatch desk, a dispatch sheet is generated; the dispatcher can dispatch the ambulance and the target hospital through the information on the dispatch sheet, and synchronize the dispatch sheet and the call record to the ambulance and the target hospital.

[0171] In one possible implementation, the dispatch record can be archived after the dispatch is completed. The dispatch record may include but is not limited to: call records, dispatch orders, dispatched ambulances, target hospitals, and timestamps, thereby ensuring the retention and reuse of information and avoiding disputes.

[0172] The method provided by the embodiment of the present invention adopts a first model to convert voice information into text information, and adopts a second model to determine the first type of entity information and the second type of entity information in the text information. The second model can focus on identifying the first address information and symptom information in the first type of entity information, and determine the actual location according to the first address information and the second address information. The first type of entity information and the second type of entity information are input into the dispatch desk, which can realize automatic filling and saving of key information of the dispatch desk, accurately obtain the key information of the rescuer, ensure the efficiency of first aid, avoid the problem of rescue delays due to repeated inquiries for key information, and enable the dispatcher to focus on dispatch work and rescue guidance more quickly.

[0173] Figure 3 A flow chart of another scheduling method provided by an application embodiment of the present invention; Figure 3 As shown, the method includes:

[0174] Step 301, 120 phone call comes in, the switch gives a voice prompt and assigns it to an idle reception dispatch desk.

[0175] Here, an idle dispatch station is determined and the distress call is forwarded to the idle dispatch station.

[0176] Step 302: The dispatcher answers the call and receives a voice message from the person in need of help.

[0177] Here, the dispatcher at the dispatch desk receives the voice message.

[0178] Step 303: The voice information is converted into text information through a voice-to-text model (equivalent to the first model).

[0179] Here, a speech-to-text model is used to convert the voice information of the person seeking help into text information.

[0180] Step 304: The text information is automatically identified as key information through a named entity recognition model (equivalent to the second model).

[0181] Here, a named entity recognition model is used to identify the key information in the text information (equivalent to the first type of entity information and the second type of entity information).

[0182] Step 305: Key information (such as address, chief complaint, age, etc.) is automatically filled in the corresponding input box of the dispatch console.

[0183] Here, it is equivalent to inputting at least the first type of entity information into the dispatching console.

[0184] Step 306: The location information (equivalent to the first address information) is input into the integrated positioning system, and the final address information (equivalent to the target address) is obtained through integrated positioning.

[0185] Here, it is equivalent to determining the actual position based on the first address information and at least one second address information, and the weight of the first address information and the weight of the second address information.

[0186] Step 307: The dispatcher dispatches an ambulance and synchronizes the on-site call record and the dispatch sheet to the ambulance and the target hospital.

[0187] Step 308: The dispatcher provides assistance at the scene.

[0188] Here, when the dispatch desk automatically determines the key information, the dispatcher can provide rescue guidance on site by phone to help the person in need (equivalent to the person asking for help or the person calling for help) to carry out initial rescue; if the rescue is successful, proceed to step 309; if the rescue is unsuccessful, proceed to step 310.

[0189] Step 309: The on-site situation is successfully resolved and the ambulance is recalled with the help-seeker's consent.

[0190] Step 310: If the situation on the scene cannot be resolved, an ambulance is sent to the scene.

[0191] Step 311, after the dispatch is completed, the dispatch record is saved in the database, including all call records, dispatched vehicles, hospitals and timestamps of each link.

[0192] Figure 4 A flowchart of another scheduling method provided by an application embodiment of the present invention; Figure 4 As shown, the method includes:

[0193] Step 401: Answer the call and receive a voice message.

[0194] Here, the dispatcher answers the call from the person seeking help and receives the voice message from the person seeking help.

[0195] Step 402: The voice information is converted into text information through a pre-trained voice-to-text model (equivalent to the first model).

[0196] Here, the scheduling system converts the received voice information into text information through a pre-trained speech-to-text model.

[0197] For example, the text message converted from voice information can be "Doctor, please come quickly, our elderly man has chest pain and sweating after using the toilet, please come and save him! The address is XX address, X building, X unit, XXX".

[0198] Step 403: The text information is identified through a pre-trained named entity recognition model (equivalent to the second model) to obtain key entity information (equivalent to the first category entity information and the second category entity information).

[0199] Here, the scheduling system can identify the text information through a pre-trained named entity recognition model to determine the key entity information in the text information.

[0200] For example, for the text information "Doctor, please come quickly, our elderly has chest pain and sweating after going to the toilet, please come and save him! The address is at XX address, X building, X unit, XXX" the named entity recognition can be "relationship with the patient: our elderly", "symptom information: chest pain, sweating" and "first address information: the address is at XX address, X building, X unit, XXX".

[0201] Among them, the first address information and the symptom information may be first-category entity information, and the relationship with the patient may be second-category entity information, and the priority of the first-category entity information is higher than that of the second-category entity information.

[0202] Step 404: The key information is automatically filled into the corresponding input box.

[0203] Here, it is equivalent to inputting at least the first type of entity information into the dispatching console.

[0204] Figure 5 A schematic diagram of a target address determination process of a scheduling method provided in an embodiment of the present invention; Figure 5 As shown, the method includes:

[0205] Step 501, entity recognition location. Here, entity recognition location refers to converting the voice information of the person in need of help into text, and then recognizing the first address information in the text information; after determining the first address information, proceed to step 506.

[0206] Step 502: GPS satellite positioning.

[0207] Here, the GPS satellite positioning position of the person in need is obtained by reading the internal positioning information of the person in need's communication device (such as a mobile phone, telephone, tablet, etc.); proceed to step 508.

[0208] Step 503: Beidou satellite positioning.

[0209] Here, the Beidou satellite positioning position of the person in need is obtained by reading the internal positioning information of the person in need's communication device; then step 509 is entered.

[0210] Step 504: Base station positioning.

[0211] Here, positioning is performed by acquiring the base station location of the operator of the communication device of the person in need; proceed to step 510 .

[0212] Step 505: Determine whether to connect to WIFI.

[0213] Determine whether the communication device of the rescuer is connected to WIFI. If it is connected to WIFI, go to step 511. If it is not connected to WIFI, go to steps 501-504.

[0214] Step 506: Match the local address database, perform error correction and similarity search.

[0215] Here, the first address information obtained is measured for similarity with the address stored in the local address database of the person in need of help, the address information with the highest matching degree is determined, and the process proceeds to step 507 .

[0216] Exemplarily, the similarity measurement may be to extract multi-dimensional features from the address information through a pre-trained feature extraction network, perform similarity calculation in a multi-dimensional space, and determine the matching degree between the first address information and each address in the address library.

[0217] Step 507: Output the address information with the maximum matching degree.

[0218] Here, the address information with the highest matching degree is output.

[0219] Step 508: Output the address information of GPS satellite positioning.

[0220] Step 509: Output the address information of Beidou satellite positioning.

[0221] Step 510: Output the address information of the base station location.

[0222] Step 511: Perform WIFI positioning.

[0223] Here, the address information is obtained through the WIFI location to which the communication device is connected.

[0224] Step 512: Output the address information of WIFI positioning.

[0225] Step 513: Voting is performed using all address information obtained.

[0226] Here, all the acquired address information (equivalent to the first address information and at least one second address information) is voted to determine the final location information (equivalent to the target address).

[0227] Step 514: Output the address information with the highest number of votes and automatically fill it into the address information column of the dispatching console.

[0228] Step 515: The dispatcher confirms.

[0229] The following provides multiple specific examples in combination with any of the above embodiments:

[0230] This proposal proposes an intelligent input method and system for emergency dispatch desk; through the relevant technology of natural language processing, the key information required in the emergency dispatch form (i.e., the first type of entity information and the second type of entity information) is automatically identified and automatically filled in and recorded. Among them, in view of the importance of location (i.e., first address information) and symptoms (symptom information) in emergency dispatch, targeted deep network recognition is performed on location and symptoms, and the location information is determined by using an integrated positioning system to obtain the location information of the caller (i.e., the target address) more timely and accurately, freeing the dispatcher's hands and allowing the dispatcher to focus more on the rescue guidance and dispatch of patients on the scene, improving emergency efficiency and competing for emergency time; the system flow of this proposal is as follows Figure 3 , Figure 4 As shown; among them,

[0231] The detailed working steps are as follows:

[0232] Step 1: When a 120 call comes in, the switch will give a voice prompt and assign it to an idle reception and dispatch desk.

[0233] Step 2: The dispatch desk accepts the call, the dispatcher answers the call, and receives the distress message sent by the caller. The system automatically fills in the key information in the dispatch form without the dispatcher having to fill it in manually. The system automatically fills in the process as follows:

[0234] After receiving the call from the person in need of help, the system converts the received voice information into text information through the pre-trained speech-to-text network (i.e., the first model), and then inputs the converted text information into the pre-trained named entity recognition network model (i.e., the second model), automatically identifies the required key information (i.e., the first type of entity information and the second type of entity information), such as the patient's location, age, gender, symptoms, etc., and automatically inputs it into the corresponding position of the dispatch desk;

[0235] Among them, the training steps of the speech-to-text network model are as follows:

[0236] 1) Obtain training data (i.e., voice training data). The training data is the voice call records of the 120 dispatch center. All call records are converted into text information as annotations. The data example is as follows:

[0237] Data voice: <<Doctor, please come quickly, our old man has chest pain and sweating after going to the toilet, please come and save him! Address is XX address X building X unit XXX>>

[0238] Label text: Doctor, please come quickly, our old man has chest pain and sweating after going to the toilet, please come and save him! The address is XX address, X building, X unit, XXX.

[0239] 2) Data enhancement. Because calls in a real 120 emergency call environment are affected by factors such as noise, distance, and volume, we enhance the training data to increase the robustness of the model (i.e., enhance the voice training data and use the enhanced voice training data to train the first model).

[0240] First, multiple data enhancement components are used on all data to alleviate the interference of environmental factors:

[0241] Volume perturbation: adjust the volume of all data. Randomly select a coefficient between [0.7, 1.4] to adjust the volume of the training data. This enhances the model's robustness to volume.

[0242] Speed ​​perturbation (i.e. speech rate perturbation): adjust the speech speed of all data. Considering that when encountering an emergency, the caller is emotional and speaks faster, so we randomly adjust the speech speed of all speech data between [0.8, 1.2] to enhance the model's robustness to speech speed;

[0243] Noise interference: Add natural noise in the range of [10dB, 500dB] to the speech of all data, such as parks, human voices, current sounds, etc. This way, the model can cover more scenarios and generalize better to noisy speech data.

[0244] Secondly, since the address and symptom information of the caller are particularly important in the 120 emergency scene, we have carried out targeted optimization on the address and symptom information (i.e. the first address information and symptom information). Through optimization, the model can be made more sensitive to address information and symptom information, and the network's recognition effect on these two types of speech can be enhanced.

[0245] In terms of data, we cut out the address and symptom parts of the speech information in the training data, and train them as a separate speech.

[0246] In terms of the loss function design of the model, based on the CTC loss function, in addition to calculating the loss function of the entire sentence, we separately annotate the address and symptoms to calculate the loss function (equivalent to determining the first loss function corresponding to the first model based on the second loss function and the third loss function), and give them a higher weight. The formula of the overall loss function is as follows:

[0247] L all =λ 1 L(S address )+λ 2 L(S symptom )+λ 3 L(S other )

[0248] In the formula, λ 1 , 2 and λ 3 Respectively represent the weights of the three loss functions (i.e., the second loss function and the third loss function). In order to make the model more sensitive to location information and symptom information, we artificially increase the weight information of these two parts (i.e., the first weight and the second weight) and reduce the weight information of other parts (i.e., the third weight). Finally, after experiments, we locate the three weights to [0.4, 0.4, 0.2] (i.e., the first weight is greater than the third weight, and the second weight is greater than the third weight). The basic loss function in the formula is the CTC loss function, and the formula is as follows:

[0249]

[0250] In the formula, the symbol S represents a training sample set (i.e., training data), which is a subset of the overall distribution. x is the output of the original data in the training sample set S after passing through the network (i.e., the prediction result of the training model). z is the label corresponding to x (i.e., the true label). p(z|x) represents the probability of restoring x to label z based on the network output x, that is, the conditional probability of z relative to x.

[0251] 3) Model training and testing. All training data are divided into training set: validation set: test set in a ratio of 6:2:2. The training set is used for model training until the model converges. The model with the best performance on the validation set and the test set is selected as the pre-trained speech-to-text model (i.e., the first model).

[0252] 4) After obtaining the pre-trained speech-to-text model, deploy the model to the system (i.e., the scheduling system). Input the voice call (i.e., voice information), and the model will output text information.

[0253] The training steps of the network model (i.e., the second model) for named entity recognition are as follows:

[0254] 1) Obtain training data. The training data is the text information of the voice call records of the 120 dispatch center, and the entities are annotated in the text sentences. The entities in this proposal are mainly divided into 4 categories, including: symptoms, address, age, gender, and relationship with the patient. Symptoms and addresses are necessary entities (i.e., the first category of entity information), and the rest are non-essential entities (i.e., the second category of entity information). The data examples are as follows:

[0255] Data text: Doctor, please come quickly, our old man has chest pain and sweating after going to the toilet, please come and save him! The address is XX address, X building, X unit, XXX.

[0256] label entity: Symptoms: chest pain, sweating;

[0257] Address: XX building X unit XXX;

[0258] Relationship with the patient: Elderly in our family --> relative;

[0259] Gender: None;

[0260] Age: None.

[0261] 2) Model training and testing.

[0262] In the design of the entity recognition model, we also focus on location information and symptom information (i.e., first address information and symptom information). The same network model, without overfitting, often performs better on simpler tasks. Therefore, we improve the general named entity recognition model, use a parallel network, and train the location and symptom information separately. The specific model structure is as follows: Figure 2 As shown; among them,

[0263] After the text is input, it passes through the BERT encoding layer to convert the text information into a vector, and then through the bidirectional LSTM (ie BI-LSTM), the context information is integrated, and then to the branch layer, where the location and symptoms (ie the first address information and symptom information) are branched. The model will better fit the location and symptom entities, and the output of the location branch only needs to judge whether the text [yes / no] is the location, thus converting a multi-classification task into a binary classification task. After the task is simple, the learning effect of the model will be greatly improved, and the same is true for the symptom branch. For other branches, it is still a multi-classification task. In this network, it is a three-classification task. The three categories are: relationship with the patient, gender, and age. This can enhance the recognition effect of the address and symptoms without incurring excessive network calculation costs. Through the above network structure, the network can more easily identify location and symptom entities and better meet the needs of emergency dispatch.

[0264] All training data are divided into training set: validation set: test set in a ratio of 6:2:2. The training set is used for model training until the model converges. The model with the best performance on the validation set and the test set is selected as the pre-trained named entity recognition model (i.e., the second model).

[0265] 3) After obtaining the pre-trained named entity recognition model, deploy the model to the system. Enter the text information, and the model will recognize the symptoms, address, age, and patient relationship information in the text (i.e., recognize the first-class entity information and the second-class entity information), and automatically fill it into the corresponding input box of the system dispatch console.

[0266] In emergency dispatch, location information is particularly critical. It is necessary to quickly and accurately obtain the location information of the caller. For example, the dispatcher may not obtain accurate location information in time, delaying the time of treatment. Therefore, in this system, the determination of location information is optimized and an integrated positioning system is used. The system flow chart is as follows Figure 5 As shown; among them,

[0267] The integrated positioning system integrates five positioning methods: deep learning entity recognition, GPS satellite positioning, Beidou satellite positioning, base station positioning, and WIFI positioning, to obtain the location information of the caller in all directions. The location information of entity recognition comes from the output of the above-mentioned entity recognition model (i.e., the second model). The location information (i.e., the first address information) obtained by deep learning entity recognition needs to be measured for similarity with the address in the local address library. The specific similarity measurement method is to use a pre-trained feature extraction network to extract multi-dimensional features. After similarity calculation in multi-dimensional space, the address with the highest matching degree is obtained. GPS satellite positioning and Beidou satellite positioning come from the internal positioning information of the mobile phone. Base station positioning is performed by collecting the base station location information of the operator. WIFI positioning obtains the location through the location of the WIFI hotspot connected to the mobile phone. Finally, the final location information is obtained through the voting mechanism. Among them, the implementation process of the voting mechanism is as follows:

[0268] According to the accuracy of each positioning method, we assign different weights to each positioning method [λ 1 ,λ 2 ,λ 3 ,λ 4 ,λ 5 ], assuming that the location information obtained by the five positioning methods are [P1, P2, P3, P4, P5] respectively, the next step is to add the weights of the same position to obtain the final position and the corresponding weight score, and finally select the position with the highest score as the final position of the patient. If the final scores are the same, the dispatcher will perform manual processing (equivalent to determining at least one second address information of the sending end of the voice information according to the preset positioning method; determining the target address according to the first address information and at least one second address information; at least inputting the first type of entity information containing the target address into the dispatch desk).

[0269] In this way, even in special circumstances where a single positioning method is insufficient or the caller cannot express clearly, the location information of the caller can be obtained in a timely and accurate manner, thus buying precious time for the dispatch of the ambulance.

[0270] Step 3: After the dispatch form is automatically filled out, the dispatcher dispatches the ambulance, and the completed dispatch form and on-site call information are synchronized to the ambulance and the target hospital;

[0271] Step 4: The dispatcher provides rescue guidance on the scene. If the situation on the scene is successfully resolved with the dispatcher's guidance, the ambulance will be recalled with the consent of the person seeking help, and the dispatch ends. If the situation cannot be resolved, the dispatch ends when the ambulance arrives at the scene.

[0272] Step 5. After the dispatch is completed, the system will retain the entire dispatch record for subsequent review, including: call records, dispatch orders, dispatch vehicles, target hospitals, and timestamps of each node.

[0273] Figure 6 A schematic diagram of a scheduling device provided by an embodiment of the present invention is shown in FIG. Figure 6 As shown, the device can be applied to intelligent electronic devices such as servers and computers; the device 60 includes: a conversion module 601, a first determination module 602, and a processing module 603;

[0274] The conversion module 601 is used to convert the received voice information into text information using the first model;

[0275] The first determination module 602 is used to determine the first type of entity information and the second type of entity information in the text information by using the second model; wherein the first type of entity information at least includes: first address information;

[0276] The processing module 603 is used to input at least the first type of entity information to the dispatching console.

[0277] Specifically, the first category of entity information at least includes: symptom information.

[0278] Specifically, the device 60 further includes: a first training module 604;

[0279] The first training module 604 is used to perform data enhancement on the speech training data, and use the data enhanced speech training data to train the first model, wherein the data enhancement includes at least one of the following: volume disturbance, speech rate disturbance and noise disturbance.

[0280] Specifically, the device 60 further includes: a second training module 605;

[0281] The second training module 605 is used to determine a second loss function corresponding to the first model according to the speech training data associated with the first type of entity information;

[0282] Determining a third loss function corresponding to the first model according to speech training data associated with the second type of entity information;

[0283] A first loss function corresponding to the first model is determined according to the second loss function and the third loss function.

[0284] Specifically, the second training module 605 is used to express the first loss function corresponding to the first model using the following formula:

[0285] Lall =λ 1 L(S address )+λ 2 L(S symptom )+λ 3 L(S other )

[0286] Among them, L(S address ) represents the second loss function corresponding to the first address information, L(S symptom ) represents the second loss function corresponding to symptom information, L(S other ) represents the third loss function corresponding to the second type of entity information, λ 1 represents the first weight of the second loss function corresponding to the first address information, λ 2 represents the second weight of the second loss function corresponding to the symptom information, λ 3 Represents a third weight of a third loss function corresponding to the second type of entity information; the first weight is greater than the third weight, and the second weight is greater than the third weight.

[0287] Specifically, the first training module 604 is used to perform a binary classification task on the text information using the second model to determine the first type of entity information in the text information;

[0288] The second model is used to perform N classification tasks on the text information to determine the second type of entity information in the text information, where N is an integer greater than or equal to 2.

[0289] Specifically, the device further includes: a second determining module 606;

[0290] Wherein, the second determination module 606 is used to determine at least one second address information of the sender of the voice information according to a preset positioning method;

[0291] determining a target address according to the first address information and the at least one second address information;

[0292] At least the first type of entity information including the target address is input into the dispatching console.

[0293] It should be noted that: when the scheduling device provided in the above embodiment implements the corresponding scheduling method, only the division of the above program modules is used as an example. In actual application, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the above-described processing. Figure 1 The embodiments of the method shown belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0294] To implement the method of the embodiment of the present invention, the embodiment of the present invention provides a scheduling device, such as Figure 7 As shown, Figure 7 A schematic diagram of the structure of another scheduling device provided by an embodiment of the present invention; the device 70 comprises: a processor 701 and a memory 702 for storing a computer program that can be run on the processor; wherein,

[0295] When the processor 701 is used to run the computer program, the processor 701 executes: using the first model to convert the received voice information into text information; using the second model to determine the first type of entity information and the second type of entity information in the text information; wherein the first type of entity information at least includes: first address information; at least inputting the first type of entity information to the dispatch station. Specifically, the first device can execute the following Figure 1 The method shown, and Figure 1 The scheduling method embodiments shown belong to the same concept, and their specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0296] In practical application, such as Figure 7 As shown, the device 70 may also include: at least one network interface 703. The various components in the scheduling device 70 are coupled together through a bus system 704. It can be understood that the bus system 704 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 704 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 7 In the figure, various buses are marked as bus system 704. There may be at least one processor 701. The network interface 703 is used for communication between the scheduling device 70 and other devices in a wired or wireless manner.

[0297] The memory 702 in the embodiment of the present invention is used to store various types of data to support the operation of the device 70 .

[0298] The method disclosed in the above embodiment of the present invention can be applied to the processor 701, or implemented by the processor 701. The processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 701 or the instruction in the form of software. The above processor 701 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 701 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 702. The processor 701 reads the information in the memory 702 and completes the steps of the above method in combination with its hardware.

[0299] In an exemplary embodiment, the scheduling device 70 can be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.

[0300] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored; when the computer program is executed by a processor, the computer program executes: using a first model to convert received voice information into text information; using a second model to determine the first type of entity information and the second type of entity information in the text information; wherein the first type of entity information at least includes: first address information; at least the first type of entity information is input to the dispatch station. Specifically, the computer program can also execute the following Figure 1 The method shown, and Figure 1 The method embodiments shown belong to the same concept, and their specific implementation processes are detailed in the method embodiments, which will not be repeated here.

[0301] In the several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0302] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0303] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0304] A person skilled in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.

[0305] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0306] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0307] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.

[0308] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A scheduling method, characterized in that: The method comprises: Converting the received voice information into text information using the first model; Using the second model, determining the first type of entity information and the second type of entity information in the text information; wherein the first type of entity information at least includes: first address information; At least the first type of entity information is input into the dispatching console.

2. The method according to claim 1, characterized in that The first category of entity information at least includes: symptom information.

3. The method according to claim 2, characterized in that The method further comprises: Performing data enhancement on the speech training data, and using the data enhanced speech training data to train the first model, wherein the data enhancement includes at least one of the following: volume disturbance, speech rate disturbance, and noise disturbance.

4. The method according to claim 2 or 3, characterized in that: The method further comprises: Determining a second loss function corresponding to the first model according to speech training data associated with the first type of entity information; Determining a third loss function corresponding to the first model according to speech training data associated with the second type of entity information; A first loss function corresponding to the first model is determined according to the second loss function and the third loss function.

5. The method according to claim 4, characterized in that The determining, according to the second loss function and the third loss function, a first loss function corresponding to the first model includes: expressing the first loss function corresponding to the first model by using the following formula: L all =λ1L(S address )+λ2L(S symptom )+λ3L(S other ) Among them, L(S address ) represents the second loss function corresponding to the first address information, L(S symptom ) represents the second loss function corresponding to symptom information, L(S other ) represents the third loss function corresponding to the second category entity information, λ1 represents the first weight of the second loss function corresponding to the first address information, λ2 represents the second weight of the second loss function corresponding to the symptom information, and λ3 represents the third weight of the third loss function corresponding to the second category entity information; the first weight is greater than the third weight, and the second weight is greater than the third weight.

6. The method according to claim 1, characterized in that The adopting the second model to determine the first type of entity information and the second type of entity information in the text information includes: Using the second model to perform a binary classification task on the text information to determine the first type of entity information in the text information; The second model is used to perform N classification tasks on the text information to determine the second type of entity information in the text information, where N is an integer greater than or equal to 2.

7. The method according to claim 1, characterized in that The method further comprises: Determine at least one second address information of a sending end of the voice information according to a preset positioning method; determining a target address according to the first address information and the at least one second address information; At least the first type of entity information including the target address is input into the dispatching console.

8. A scheduling device, characterized in that: The device comprises: a conversion module, a first determination module, and a processing module; wherein the conversion module is used to convert the received voice information into text information using the first model; The first determination module is used to determine the first type of entity information and the second type of entity information in the text information by using the second model; wherein the first type of entity information at least includes: first address information; The processing module is used to input at least the first type of entity information into the dispatching console.

9. A scheduling device, characterized in that: The device includes a network interface, a memory and a processor; wherein the network interface is used to realize connection communication between components; the memory is used to store a computer program that can be run on the processor; and the processor is used to execute the method described in any one of claims 1 to 7 when running the computer program.

10. A computer storage medium, characterized in that: The computer storage medium stores a computer program, and when the computer program is executed by at least one processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Emergency team real-time voice instruction analysis and collaborative optimization method and system

    CN120783747A