Address identification method, address identification device, electronic equipment and storage medium
By combining the trained address recognition model and the large language model, the address recognition and verification of dialogue data is solved, and the problem of inaccurate address information recognition in dialogue data is achieved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510226387.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately identify user address information in dialogue data, and is greatly affected by oral expression problems and pronunciation escaping problems.
Two models with different characteristics are used for fused address recognition. The first model is a trained address recognition model, dedicated to address recognition tasks; the second model is based on a large language model and has strong language processing capabilities. When the verification result is not passed, the dialogue text is corrected to improve the accuracy of address recognition.
Through model fusion and correction processing, the identification accuracy of address information in dialogue text is significantly improved, and the problem of identification errors and fuzzy in the prior art is solved.
Smart Images

Figure CN120068880A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to an address recognition method, an address recognition device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In actual business scenarios, it is often necessary to extract the address information of a user from the conversation data between a customer service and the user. Affected by factors such as the user's oral expression problems and speech transference problems, the address information in the conversation data may have semantic errors, ambiguity, etc., resulting in inaccurate recognition of the address information. Summary of the Invention
[0003] The present disclosure provides an address recognition method, a device, an electronic device, and a computer-readable storage medium to improve the accuracy of address recognition for conversation texts.
[0004] In a first aspect, the present disclosure provides an address recognition method, which includes:
[0005] Inputting a target conversation text into a first model, and extracting target address information corresponding to the target conversation text according to the first model, where the first model is a trained address recognition model;
[0006] Inputting the target address information and the target conversation text into a second model, and verifying the target address information according to the second model to obtain a verification result of the target address information, where the second model is determined according to a large language model;
[0007] In the case where the verification result is that the verification fails, performing a correction process on the target conversation text according to the second model so that the first model extracts corrected address information according to the corrected target conversation text; and generating an address recognition result corresponding to the target conversation text according to the corrected address information.
[0008] In a second aspect, the present disclosure provides an address recognition device, which includes:
[0009] An extraction module, configured to input a target conversation text into a first model, and extract target address information corresponding to the target conversation text according to the first model, where the first model is a trained address recognition model;
[0010] A verification module, configured to input the target address information and the target conversation text into a second model, and verify the target address information according to the second model to obtain a verification result of the target address information, where the second model is determined according to a large language model;
[0011] A correction module, configured to, when the verification result is that the verification fails, correct the target dialogue text according to the second model, so that the first model extracts corrected address information according to the corrected target dialogue text; and generate an address recognition result corresponding to the target dialogue text according to the corrected address information.
[0012] In a third aspect, the present disclosure provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor, so that the at least one processor can execute the above address recognition method.
[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, wherein the computer program realizes the above address recognition method when executed by a processor.
[0014] The address recognition method, address recognition device, electronic device, and storage medium provided by the embodiments of the present disclosure can perform fused address recognition based on two models with different characteristics. On the one hand, the first model is a trained address recognition model, that is, it is a model dedicated to the address recognition task. Therefore, according to the first model, the target address information corresponding to the target dialogue text can be accurately extracted. On the other hand, the second model is determined according to the large language model, so it has strong language processing capabilities. Thus, according to the second model, the target address information extracted by the first model can be accurately verified, and when the verification fails, the target dialogue text is corrected, so that the first model further extracts corrected address information according to the corrected target dialogue text, and then generates an address recognition result corresponding to the target dialogue text according to the corrected address information, thereby effectively improving the accuracy of address recognition.
[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0016] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. They are used to explain the present disclosure together with the embodiments of the present disclosure, and do not constitute a limitation to the present disclosure. By describing the detailed exemplary embodiments with reference to the drawings, the above and other features and advantages will become more obvious to those skilled in the art. In the drawings:
[0017] Figure 1Application scenario diagram of the address recognition method and device provided by the embodiments of the present disclosure;
[0018] Figure 2 Flowchart of an address recognition method provided by the embodiments of the present disclosure;
[0019] Figure 3 Schematic diagram of extracting corrected address information in the embodiments of the present disclosure;
[0020] Figure 4 Schematic diagram of an address recognition method provided by the embodiments of the present disclosure;
[0021] Figure 5 Block diagram of an address recognition device provided by the embodiments of the present disclosure;
[0022] Figure 6 Block diagram of an electronic device provided by the embodiments of the present disclosure. Detailed implementation manners
[0023] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.
[0024] Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0025] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms "include" and / or "consist of" are used in this specification, the specified features, wholes, steps, operations, elements, and / or components are present, but one or more other features, wholes, steps, operations, elements, components, and / or their groups are not excluded. "Connection" or "coupling" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0027] Unless otherwise defined, all terms (including technical and scientific terms) used herein shall have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.
[0028] In the technical solutions of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. The use of user data in this technical solution follows relevant national laws and regulations (for example, "Information Security Technology - Personal Information Security Specification", etc.). For example, corresponding regulatory measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; when using personal information, the clear identity indication is eliminated to avoid precisely locating a specific individual.
[0029] In actual application scenarios such as the telemarketing scenario, it is usually necessary to obtain or confirm the user's address information from the conversation data. Affected by factors such as the user's oral expression problems and voice transference problems, the existing technical solutions often cannot achieve accurate recognition of the address information.
[0030] In view of this, the embodiments of the present disclosure provide an address recognition method, an address recognition device, an electronic device, and a storage medium, which can perform fused address recognition based on two models with different characteristics. On the one hand, the first model is a trained address recognition model, that is, it is a model dedicated to the address recognition task. Therefore, according to the first model, the target address information corresponding to the target conversation text can be accurately extracted. On the other hand, the second model is determined according to the large language model, so it has strong language processing capabilities. Thus, according to the second model, the target address information extracted by the first model can be accurately verified, and in the case of failed verification, the target conversation text is corrected so that the first model further extracts the corrected address information according to the corrected target conversation text. Then, according to the corrected address information, the address recognition result corresponding to the target conversation text is generated, thereby effectively improving the accuracy of address recognition.
[0031] Figure 1 It is an application scenario diagram of the address recognition method and device provided by the embodiments of the present disclosure.
[0032] As Figure 1As shown in the figure, the application scenario of the embodiments of the present disclosure may include a terminal device 101, a network 103, and a server 102. The network 103 is used to provide a medium for a communication link between the terminal device 101 and the server 102. The network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0033] A user may use the terminal device 101 to interact with the server 102 through the network 103 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0034] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, etc.
[0035] The server 102 may be a server providing various services, such as a background management server that supports the websites browsed by users using the terminal device 101 (only as an example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0036] It should be noted that the address recognition method and device provided by the embodiments of the present disclosure may be executed by the server 102. Correspondingly, the address recognition method and device provided by the embodiments of the present disclosure may be set in the server 102. The address recognition method and device provided by the embodiments of the present disclosure may also be executed by a server or a server cluster different from the server 102 and capable of communicating with the terminal device 101 and / or the server 102. Correspondingly, the address recognition method and device provided by the embodiments of the present disclosure may also be set in a server or a server cluster different from the server 102 and capable of communicating with the terminal device 101 and / or the server 102.
[0037] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in are merely illustrative. According to actual needs, there may be any number of terminal devices, networks, and servers.
[0038] Figure 2 is a flowchart of an address recognition method provided by the embodiments of the present disclosure. Referring to Figure 2 , the method includes:
[0039] Step S210: Input the target dialogue text into a first model, and extract the target address information corresponding to the target dialogue text according to the first model, where the first model is a trained address recognition model.
[0040] Among them, the target dialogue text refers to the dialogue text containing address information, that is, the dialogue text related to the address. Specifically, the target dialogue text can include one or more rounds of dialogue texts.
[0041] Among them, the dialogue text generated during a round of question-and-answer interaction among multiple call objects is a round of dialogue text. For example, call object A asks call object B about the address "May I ask where is your current residence?", and call object B feedbacks the address "In a certain district of a certain city". The above dialogue text is a round of dialogue text.
[0042] In an actual scenario, it may be necessary to complete the address inquiry through multiple rounds of question-and-answer interaction among call objects. At this time, the target dialogue text can include multiple rounds of dialogue texts generated by multiple rounds of question-and-answer interaction.
[0043] Among them, the first model refers to the trained address recognition model. The address recognition model can be any model used for address recognition tasks. For example, a conditional random field model, a convolutional neural network model, a long short-term memory network model, etc. The embodiments of the present disclosure do not limit this. By training the address recognition model, the first model can be obtained.
[0044] In an optional implementation manner, the address recognition model can be trained based on training dialogue samples composed of address-related words and non-address-related words, so as to obtain the trained address recognition model, that is, the first model.
[0045] Among them, the target address information refers to the address information obtained by the first model for address recognition of the target dialogue text. For example, for the target dialogue text "The address is in District B, City A", the first model can extract the target address information as "District B, City A".
[0046] It should be noted that in order to improve the data integrity of the target address information, the target address information can include address words and the administrative region types corresponding to the address words. Among them, the administrative region type corresponding to the address word is used to represent the administrative region level corresponding to the address word, such as provincial administrative region, municipal administrative region, etc. Exemplarily, the target address information corresponding to the above target dialogue text can also be "City: City A, District: District B". The embodiments of the present disclosure do not limit this.
[0047] It should also be noted that in the case where the target dialogue text includes multiple rounds of dialogue texts, the target address information can be determined according to multiple address information extracted by the first model from the multiple rounds of dialogue texts.
[0048] For example, the target dialogue text includes: the first-round dialogue text "Agent: May I ask where your house is located? User: In City AA."; the second-round dialogue text "Agent: Could you please tell me the detailed address? User: The specific location is in District XX." Correspondingly, the target address information corresponding to the target dialogue text extracted by the first model can be "District XX, City AA".
[0049] Step S220: Input the target address information and the target dialogue text into the second model, and verify the target address information according to the second model to obtain the verification result of the target address information, where the second model is determined according to the large language model.
[0050] Among them, the large language model is a pre-trained model, which has natural language inference technology and can understand, generate, and reason about natural language.
[0051] Specifically, the second model can be determined according to any existing large language model. For example, directly use the large language models GPT and Qwen as the second model, or fine-tune the model parameters of the large language model according to the preset dialogue text library to obtain the second model. The embodiments of the present disclosure do not limit this.
[0052] It should be noted that the prompt information corresponding to the address information verification task can be generated according to the target address information and the target dialogue text, so as to input the prompt information into the second model, so that the second model outputs the verification result of the target address information.
[0053] Exemplarily, the prompt information prompt_verify corresponding to the address information verification task can be:
[0054] prompt_verify =
[0055] "Task: Please verify the identified address information according to the content of the dialogue text.
[0056] Dialogue text:
[0057] Agent: Which district and county in Beijing is your house in? User: No, in Tianjing.
[0058] Agent: Where did you say? User: Binhai Binhai New Area.
[0059] Address information: [{"name": "Address", "value": "Tianjing"},
[0060] {"name": "Address", "value": "Binhai Binhai New Area"}]."
[0061] Correspondingly, the second model verifies the target address information according to the prompt information and outputs the verification result of the target address information, such as verification passed or verification failed.
[0062] In an alternative implementation, in order to improve the verification accuracy of the second model for the target address information, the preset verification type, the target address information, and the target dialogue text can also be jointly input into the second model.
[0063] Among them, the preset verification type is used to represent the verification factors of the second model for the target address information. The preset verification type can include word order coherence verification, semantic accuracy verification, and / or address type verification.
[0064] Among them, the word order coherence verification is used to verify whether the target address information has coherent word order. For example, whether the administrative region sorting of the target address information is confused, such as the target address information being "a certain county, a certain city".
[0065] Among them, the semantic accuracy verification is used to verify whether there are address words with semantic errors in the target address information. The address type verification is used to verify whether the target address information is of the address type.
[0066] Step S230: In the case where the verification result is verification failed, the target dialogue text is corrected according to the second model so that the first model extracts the corrected address information based on the corrected target dialogue text; according to the corrected address information, an address recognition result corresponding to the target dialogue text is generated.
[0067] Among them, in order to improve the verification efficiency of the second model for the target address information and the correction efficiency of the target dialogue text, the address information verification task and the dialogue text correction task can be combined into the same prompt information, so that when the second model determines that the verification result of the target address information is verification failed, the corrected target dialogue text is output.
[0068] Exemplarily, the prompt information prompt_verify corresponding to the address information verification task and the dialogue text correction task can be:
[0069] prompt_verify =
[0070] "Task: Please verify or correct the recognized address information according to the content of the dialogue text.
[0071] Dialogue text:
[0072] Agent: Which district or county in Beijing is your house in? User: No, it's in Tianjing.
[0073] Agent: Where did you say? User: Binhai Binhai New Area.
[0074] Address information: [{"name": "Address", "value": "Tianjing"},
[0075] {"name": "Address", "value": "Binhai New Area"}
[0076] Requirements:
[0077] 1) Output the discrimination result in JSON structure, where the JSON primary keys are the if_correct field and the corrected_text field.
[0078] 2) If the recognition is correct, the value of if_correct is "Verification Passed" and the value of corrected_text is an empty string;
[0079] 3) If the recognition is incorrect, the value of if_correct is "Verification Failed" and the value of corrected_text is the corrected dialogue text.
[0080] It should be noted that the second model can perform multiple correction processes on the target dialogue text. That is, after the second model corrects the target dialogue text, the first model needs to continue to extract the address information based on the corrected target dialogue text, and the second model needs to verify the address information. Thus, in the case of verification failure, the second model continues to correct the target dialogue text.
[0081] Correspondingly, the corrected address information is the address information extracted by the first model from the corrected target dialogue text corresponding to the last correction process. Since the corrected address information is extracted from the corrected target dialogue text, compared with the target address information, the corrected address information has higher content accuracy. Thus, based on the corrected address information, the address recognition result corresponding to the target dialogue text can be obtained.
[0082] In an optional implementation, when the verification result is "Verification Passed", generate the address recognition result corresponding to the target dialogue text according to the target address information.
[0083] Since the verification result of the target address information is "Verification Passed", it indicates that the target address information has high accuracy. Therefore, the target address information can be directly used as the address recognition result corresponding to the target dialogue text.
[0084] In the embodiments of the present disclosure, address recognition can be performed by fusing two models with different characteristics. On the one hand, the first model is a trained address recognition model, that is, it is a model dedicated to the address recognition task. Therefore, according to the first model, the target address information corresponding to the target dialogue text can be accurately extracted. On the other hand, the second model is determined according to the large language model, so it has strong language processing capabilities. Thus, according to the second model, the target address information extracted by the first model can be accurately verified, and in the case where the verification fails, the target dialogue text is corrected so that the first model further extracts the corrected address information according to the corrected target dialogue text. Then, according to the corrected address information, the address recognition result corresponding to the target dialogue text is generated, thereby effectively improving the accuracy of address recognition.
[0085] In an alternative implementation, in order to improve the accuracy of the corrected address information, the target dialogue text can be corrected based on the second model and the first model can perform iterative extraction of address information from the corrected target dialogue text to obtain the corrected address information.
[0086] For ease of understanding, Figure 3 a flowchart for extracting the corrected address information in the embodiments of the present disclosure is shown. Referring to Figure 3 , in step S230, the target dialogue text is corrected according to the second model so that the first model extracts the corrected address information according to the corrected target dialogue text, which can be implemented in the following manner:
[0087] Step S310: Perform the (i + 1)-th correction process on the target dialogue text according to the second model to obtain the (i + 1)-th corrected target dialogue text.
[0088] Wherein, i is a natural number. When i is 0, it indicates that the second model performs the first correction process on the target dialogue text, that is, the trigger condition for this step at this time is that the verification result of the target address information fails the verification. When i is greater than 0, it indicates that the second model performs a non-first correction process on the target dialogue text, and then the trigger condition for this step is that the verification result of the i-th address information corresponding to the i-th corrected target dialogue text does not meet the preset conditions.
[0089] Step S320: Input the (i + 1)-th corrected target dialogue text into the first model, and extract the (i + 1)-th address information corresponding to the (i + 1)-th corrected target dialogue text according to the first model.
[0090] Step S330: Input the (i + 1)-th address information corresponding to the (i + 1)-th corrected target dialogue text and the (i + 1)-th corrected target dialogue text into the second model to obtain the verification result of the (i + 1)-th address information corresponding to the (i + 1)-th corrected target dialogue text.
[0091] Step S340: Determine whether the verification result meets a preset condition.
[0092] Among them, the preset condition includes: the verification result is verification passed and / or the number of correction processes reaches a preset number threshold.
[0093] It should be noted that the preset number threshold can be adaptively set according to actual application requirements, and the embodiments of the present disclosure do not limit this. By recording and counting the number of correction processes, after each correction process, the value of the number of correction processes is incremented by 1. When the number of correction processes reaches the preset number threshold and / or the verification is passed, it is determined that the verification result meets the preset condition, thus effectively avoiding the second model from over-correcting the target dialogue text.
[0094] Step S350: When the verification result of the (i + 1)-th address information corresponding to the target dialogue text after the (i + 1)-th correction meets the preset condition, obtain the corrected address information according to the (i + 1)-th address information corresponding to the target dialogue text after the (i + 1)-th correction.
[0095] It should be noted that if the verification result of the (i + 1)-th address information does not meet the preset condition, the value of i is incremented by 1 and steps S310 to S350 are continued to be executed.
[0096] For example, perform the (i + 2)-th correction process on the target dialogue text according to the second model to obtain the target dialogue text after the (i + 2)-th correction, extract the (i + 2)-th address information corresponding to the target dialogue text after the (i + 2)-th correction according to the first model, and obtain the verification result of the (i + 2)-th address information through the second model. Thus, when the verification result meets the preset condition, obtain the corrected address information according to the (i + 2)-th address information. When the verification result does not meet the preset condition, continue to perform the (i + 3)-th correction process on the target dialogue text according to the second model.
[0097] In the embodiments of the present disclosure, when the verification result of the address information extracted by the first model by the second model does not meet the preset condition, the second model performs repeated correction processes on the target dialogue text so that the first model can extract address information according to the corrected target dialogue text, thereby improving the information accuracy of the corrected address information.
[0098] In an alternative implementation, to improve the correction effect of the second model on the target dialogue text, the second model can be instructed to correct the target dialogue text through a preset correction strategy. Accordingly, the (i + 1)-th correction process of the target dialogue text by the second model includes: inputting the preset correction strategy into the second model so that the second model performs the (i + 1)-th correction process on the target dialogue text according to the preset correction strategy, where the preset correction strategy includes at least one of a dialogue merging strategy, a word order adjustment strategy, and a semantic adjustment strategy; in the case where the preset correction strategy is the dialogue merging strategy, the second model performing the (i + 1)-th correction process on the target dialogue text according to the preset correction strategy includes: when the target dialogue text includes multi-round dialogue texts with a context relationship, integrating the multi-round dialogue texts into one round of dialogue text according to the second model; in the case where the preset correction strategy is the word order adjustment strategy, the second model performing the (i + 1)-th correction process on the target dialogue text according to the preset correction strategy includes: arranging the address information included in the target dialogue text in the order of the preset administrative region levels according to the second model; in the case where the preset correction strategy is the semantic adjustment strategy, the second model performing the (i + 1)-th correction process on the target dialogue text according to the preset correction strategy includes: semantically adjusting the address information with semantic deviation included in the target dialogue text according to the second model.
[0099] Among them, the multi-round dialogue texts with a context relationship refer to that the multi-round dialogue texts correspond to the same address inquiry intention. For example, when the seat asks the user for address information, multiple rounds of conversations are carried out between the seat and the user, so the multi-round dialogue texts corresponding to the multiple rounds of conversations have a context relationship.
[0100] Specifically, when integrating the multi-round dialogue texts into one round of dialogue text, the discrete information included in the multi-round dialogue texts can be merged to obtain one round of dialogue text. For example, the provincial and municipal address information of the user is included in the first-round dialogue text, while the district and county address information of the user is included in the second-round dialogue text. Therefore, the provincial and municipal address information and the district and county address information can be merged.
[0101] Among them, the word order adjustment strategy is used to instruct the second model to arrange the address information included in the target dialogue text in the order of the preset administrative region levels. For example, the preset administrative region level order can be provincial administrative region, municipal administrative region, and district and county administrative region. Accordingly, in the target dialogue text corrected by the second model, the included address information can be XX District, XX City, XX Province.
[0102] Among them, the semantic adjustment strategy is used to instruct the second model to semantically adjust the address information with semantic deviation included in the address information.
[0103] For example, due to ASR escape errors, the text information representation of some words in the target dialogue text is inaccurate. For example, "Tianjin" is escaped as "Tianjing". Another example is that due to the user's own language expression, such as tone pauses, there is duplicate address information in the target dialogue text, such as "Tianjin Tianjin". The second model can semantically adjust the address information with semantic deviation in the target dialogue text to improve the accuracy of the text content of the target dialogue text.
[0104] It should be noted that when the second model corrects the target dialogue text, the above preset correction strategies can be used together.
[0105] Exemplarily, for the multi-round dialogue text "Agent: Which district or county in Beijing is your house in? User: No, in Tianjing", "Agent: Where do you mean? User: Binhai Binhai New Area", the corrected target dialogue text obtained after the second model corrects according to the above preset correction strategies can be "Agent: Which district or county in Beijing is your house in? User: No, in Binhai, Tianjin, Binhai New Area".
[0106] In the embodiment of the present disclosure, by inputting the preset correction strategy into the second model, the second model integrates the context of the target dialogue text according to the preset correction strategy, integrates the discrete information of the multi-round dialogue text into a single-round dialogue text, and corrects the word order and semantic errors in the target dialogue text, thereby improving the correction effect of the target dialogue text, so that the first model can extract more accurate address information according to the corrected target dialogue text.
[0107] In an optional implementation manner, a large amount of dialogue text may be generated between multiple call objects, and only some of the dialogue text contains address information. Therefore, in order to improve the extraction efficiency and accuracy of the target dialogue text, the second model can be used to identify the dialogue text to determine whether the dialogue text contains address information.
[0108] Correspondingly, the target dialogue text is obtained in the following manner: for any round of dialogue text in the multi-round historical dialogue text, according to the second model, identify whether any round of dialogue text is the target dialogue text containing address information, and obtain M initial target dialogue texts, where M is a natural number; according to the second model, determine whether there is a context relationship between the M initial target dialogue texts, and splice at least two initial target dialogue texts with a context relationship to obtain N processed target dialogue texts, where N is a natural number less than or equal to M.
[0109] Among them, the prompt information corresponding to the address association recognition task can be constructed based on the historical dialogue text, and then the prompt information is input into the second model, so that the second model can determine whether the dialogue text is the target dialogue text containing address information.
[0110] For example, the prompt word prompt_assess_i corresponding to the address association recognition task constructed for the i-th round of dialogue text in the historical dialogue text is as follows:
[0111] prompt_assess_i =
[0112] "Task: Please determine whether the current dialogue is related to an address according to the dialogue content.
[0113] Dialogue: Agent: Where do you mean? User: Binhai Binhai New Area.
[0114] Output the judgment result: Yes or No (Yes means related, No means not related)", so as to obtain the output result output_i of the second model. For example, output_i is "Yes".
[0115] Furthermore, when the number of initial target dialogue texts is multiple, that is, when M>1, the dialogue rounds corresponding to the M initial target dialogue texts can be determined. Thus, when the difference between the dialogue rounds corresponding to multiple target dialogue texts among the M initial target dialogue texts is within a preset range, the multiple target dialogue texts are input into the second model so that the second model can determine whether the multiple target dialogue texts have a context relationship.
[0116] Among them, having a context relationship means that multiple target dialogue texts correspond to the same address consultation intention. Exemplarily, the second-round and third-round initial target dialogue texts are input into the second model, so that the second model can determine whether the above initial target dialogue texts have a context relationship. Thus, when the second model determines that the above initial target dialogue texts have a context relationship, the above initial target dialogue texts are concatenated into a target dialogue text, so as to obtain N processed target dialogue texts.
[0117] For example, the second-round initial target dialogue text text_2 is "Agent: Which district or county in Beijing is your house in? User: No, in Tianjin.", and the third-round initial target dialogue text text_3 is "Agent: Where do you mean? User: Binhai Binhai New Area." By concatenating the above initial target dialogue texts, the processed target dialogue text "Agent: Which district or county in Beijing is your house in? User: No, in Tianjin. Agent: Where do you mean? User: Binhai Binhai New Area." can be obtained.
[0118] In the embodiments of the present disclosure, by using a second model (such as Qwen, GPT, etc.) to classify the dialogue text, it is determined whether the text contains address-related information, so as to improve the accuracy of address relevance judgment. By the output of the model being yes / no, the response efficiency is also effectively improved. Moreover, in the embodiments of the present disclosure, the second model is also used to determine whether the initial target dialogue text has a context relationship, so as to splice the initial target dialogue texts with context relationships, thereby improving the information integrity of the target dialogue text.
[0119] In an alternative implementation, in order to improve the integrity of the target address information extracted by the first model, a preset address vocabulary can be set in the first model, so that the first model can perform structured processing on the extracted address words according to the preset address vocabulary.
[0120] Correspondingly, extracting the target address information corresponding to the target dialogue text according to the first model includes: extracting the address text words corresponding to the target dialogue text according to the first model; matching the address text words with a plurality of preset address words included in the preset address vocabulary, and determining the administrative region type corresponding to the address text words according to the matching result, where the preset address vocabulary includes the administrative region type corresponding to the preset address words and the hierarchical relationship to which the preset address words belong; when the administrative region type corresponding to the address text words is a preset type, determining the hierarchical relationship to which the address text words belong and the filled address words corresponding to the hierarchical relationship according to the hierarchical relationship to which the preset address words belong; and performing filling processing on the address text words according to the filled address words to obtain the target address information.
[0121] Among them, when matching the address text words with a plurality of preset address words included in the preset address vocabulary, the semantic similarity between the address text words and the preset address words can be calculated to obtain the matching result, and the matching result may include the preset address words that match the address text words.
[0122] Among them, the hierarchical relationship of the preset address words can be divided according to the administrative region type. For example, provincial administrative regions, municipal administrative regions, and district / county-level administrative regions are the first level, the second level, and the third level respectively.
[0123] Correspondingly, the first level is the upper level of the second level, and the second level is the upper level of the third level. Exemplarily, the preset address word corresponding to the first level can be "Province A", the preset address word corresponding to the second level can be the municipal administrative region corresponding to Province A, such as "City B", and the preset address word corresponding to the third level can be the district / county-level administrative region corresponding to "City B", such as "County C".
[0124] Thus, for the above preset address word "County C", the hierarchical relationship to which it belongs can be "First level: Province A, Second level: City B, Third level: County C".
[0125] Among them, the preset type can be any administrative region type other than the administrative region type corresponding to the first level in the hierarchical relationship. Specifically, the preset type can be adaptively set according to actual needs. For example, the preset type is a municipal administrative region type, a district and county administrative region type, etc. The embodiments of the present disclosure do not limit this.
[0126] Thus, when the administrative region type corresponding to the address text word is the preset type, the hierarchical relationship to which the address text word belongs can be determined according to the hierarchical relationship to which the preset address word matching the address text word belongs, so as to obtain the filled address word corresponding to this hierarchical relationship.
[0127] Among them, the filled address word can be determined according to the preset address word of the address text word in the previous level in the hierarchical relationship to which the address text word belongs. For example, when the address text word is "C County", the hierarchical relationship to which it belongs is "First level: A Province, Second level: B City, Third level: C County". Correspondingly, the filled address words corresponding to it can be "A Province" and "B City".
[0128] Correspondingly, the address text word can be filled according to the filled address word to obtain the target address information, such as "Province: A Province, City: B City, District / County: C County".
[0129] It should also be noted that the target dialogue text may contain address text words corresponding to multiple administrative region types. Therefore, the filled address words can be further de-duplicated with all the extracted address text words to obtain the final filled address words.
[0130] For example, when the preset type is the district and county administrative region type, for all address text words, when the administrative region type corresponding to any address text word is the district and county administrative region type, it is further determined whether the remaining address text words corresponding to the provincial and municipal administrative region types are extracted. If not, the filled address word corresponding to the address text word is obtained based on the hierarchical relationship to which the address text word belongs. If so, the target address information is obtained according to the order of the address text words corresponding to the three-level administrative region types of province, city, and district / county.
[0131] In an alternative implementation, since there may be a situation where the address information in the target dialogue text is not standardized, resulting in address words with deviations in the target address information extracted by the first model from the target dialogue text, which affects the accuracy of the target address information. In order to improve the address information extraction effect of the first model for the target dialogue text with semantic deviations, the address recognition model can be trained according to the deviation dialogue samples corresponding to the preset semantic deviation types, so as to obtain the first model.
[0132] Correspondingly, the first model is trained as follows: perform semantic deviation processing on the standard address words included in the preset dialogue samples according to the preset semantic deviation types to obtain deviation dialogue samples corresponding to the preset dialogue samples, where the preset semantic deviation types at least include semantic repetition type, semantic pause type, and / or dialect conversion type; construct training samples based on the preset dialogue samples and the deviation dialogue samples, input the training samples into the initial address recognition model to obtain the predicted address words corresponding to the training samples and the predicted semantic deviation types corresponding to the training samples; compare the standard address words corresponding to the training samples with the predicted address words to obtain a first comparison result, and compare the preset semantic deviation types corresponding to the training samples with the predicted semantic deviation types to obtain a second comparison result; adjust the training parameters of the initial address recognition model according to the first comparison result and the second comparison result to obtain the first model.
[0133] Among them, the preset dialogue samples can be obtained from the dialogue sample library. The preset dialogue samples contain standard address words, that is, address words without semantic deviation.
[0134] Among them, when the preset semantic deviation type is the semantic repetition type, by performing semantic repetition processing on the preset dialogue samples, the generated deviation dialogue samples contain repeated address words, such as "Beijing Beijing".
[0135] Among them, when the preset semantic deviation type is the semantic pause type, by performing semantic pause processing on the preset dialogue samples, a deviation dialogue sample with semantic pauses is generated. For example, insert a preset punctuation mark into the standard address word "Beijing City" in the preset dialogue sample to obtain the address word with semantic pauses "North, Beijing City".
[0136] Among them, when the preset semantic deviation type is the dialect conversion type, by searching for the dialect conversion words corresponding to the standard address words in the preset dialogue samples (such as words with similar pronunciations to the standard address words, dialect words corresponding to the standard address words), the standard address words are replaced with dialect conversion words.
[0137] Thus, by comparing the standard address words with the predicted address words, for example, calculating the semantic similarity between the words, a first comparison result is obtained. Comparing the preset semantic deviation types corresponding to the training samples with the predicted semantic deviation types, for example, calculating the prediction accuracy of the initial address recognition model for the semantic deviation types, a second comparison result is obtained. Then, adjust the training parameters of the initial address recognition model according to the first comparison result and the second comparison result to improve the prediction accuracy of the model for address words and semantic deviation types, and obtain the trained address recognition model, that is, the first model.
[0138] It should be noted that in order for the first model to effectively identify address words of semantic repetition type, semantic pause type, and / or dialect transformation type, the character features of the training samples and the position encoding features corresponding to the training samples can be combined during the identification.
[0139] Correspondingly, input the training samples into the initial address recognition model to obtain the predicted address words corresponding to the training samples and the predicted semantic deviation types corresponding to the training samples, including: input the training samples into the feature extraction module of the initial address recognition model to obtain the character encoding vectors of the characters in the training samples and the position encodings corresponding to the characters; input the character encoding vectors and the position encodings into the address extraction module of the initial address recognition model to obtain the position intervals corresponding to the candidate address words; according to the position intervals corresponding to the candidate address words, search for the word encoding vectors corresponding to the candidate address words in the character encoding vectors, and input the word encoding vectors into the semantic deviation classification module of the initial address recognition model to obtain the predicted semantic deviation types corresponding to the candidate address words; input the predicted semantic deviation types and the candidate address words into the address processing module of the initial address recognition model to obtain the predicted address words corresponding to the training samples.
[0140] Exemplarily, assume that the training sample is "The address is Tian, Jin City". By inputting the training sample into the feature extraction module, the character encoding vectors of each character in the training sample and the position encodings corresponding to the characters (i.e., position indices, such as 0, 1) can be obtained. Correspondingly, input the character encoding vectors and the position encodings into the address extraction module to obtain the position intervals corresponding to the candidate address words. For example, the position interval of the candidate address word "Tian," is (3, 4), and the position interval of "Jin" is 5. Then, according to the position intervals, search for the word encoding vectors corresponding to the candidate address words in the character encoding vectors, and further input the word encoding vectors into the semantic deviation classification module to obtain the predicted semantic deviation types corresponding to the candidate address words, such as the semantic pause type. Finally, input the predicted semantic deviation types and the candidate address words into the address processing module to obtain the predicted address words corresponding to the training sample, such as "Tianjin".
[0141] Furthermore, corresponding address processing rules can be configured for the address processing module. For example, when the predicted semantic deviation type corresponding to the candidate address word is the semantic pause type, remove the punctuation marks in the candidate address word and / or merge adjacent candidate address words. When the predicted semantic deviation type corresponding to the candidate address word is the semantic repetition type, merge the repeated character pairs at adjacent positions in the candidate address word. When the predicted semantic deviation type corresponding to the candidate address word is the dialect transformation type, query the word mapping table of dialect words and standard words to obtain the processed candidate address word.
[0142] In the embodiments of the present disclosure, when performing address word prediction, the model not only combines the character features of the training samples, but also combines the position encoding features corresponding to the training samples, thereby improving the model's context analysis ability for each character in the training samples, improving the prediction effect of address words, and the prediction effect of semantic deviation types.
[0143] In an alternative implementation, in order to improve the semantic deviation classification effect of the semantic deviation classification module on the word encoding vector of the candidate address word, the administrative level label corresponding to the candidate address word and the label anomaly probability corresponding to the candidate address word can also be jointly input into the semantic deviation classification module.
[0144] Correspondingly, inputting the word encoding vector into the semantic deviation classification module to obtain the predicted semantic deviation type corresponding to the candidate address word includes: inputting the word encoding vector and the position interval corresponding to the candidate address word into the label prediction module of the initial address recognition model to obtain the administrative level label corresponding to the candidate address word; inputting the word encoding vector and the symbol type label corresponding to the candidate address word into the punctuation anomaly recognition module of the initial address recognition model to obtain the punctuation anomaly probability corresponding to the candidate address word; inputting the administrative level label, the label anomaly probability, and the word encoding vector into the semantic deviation classification module to obtain the predicted semantic deviation type corresponding to the candidate address word.
[0145] Among them, the administrative level label is used to identify the administrative region type corresponding to the candidate address word. For example, the administrative level label "Province" is used to identify the provincial administrative region type.
[0146] Among them, the symbol type label corresponding to the candidate address word is used to identify the symbol type corresponding to each character in the candidate address word, such as Chinese characters, numbers, punctuation marks, etc. Specifically, before inputting the training sample into the initial address recognition model, the corresponding symbol type label can be generated for each character in the training sample.
[0147] Among them, the punctuation anomaly probability is used to characterize the probability of abnormal punctuation marks appearing in the candidate address word, such as the probability of the candidate address word containing punctuation marks that cause semantic pauses.
[0148] Thus, by jointly inputting the administrative level label, the punctuation anomaly probability, and the word encoding vector into the semantic deviation classification module, the semantic deviation classification module can determine the probability that the candidate address word belongs to each semantic deviation type to obtain the predicted semantic deviation type.
[0149] Specifically, a standard address word vector library can be built into the semantic deviation classification module to improve the classification effect of the semantic deviation classification module on candidate address words of the dialect conversion type.
[0150] It should be noted that the specific structures of the modules in the initial address recognition model in the present disclosure can be adaptively set according to actual needs, and the embodiments of the present disclosure do not limit this. For example, the feature extraction module can be composed of a convolutional layer, a GRU layer, etc., and the punctuation anomaly recognition module can be composed of a convolutional layer, a pooling layer, a linear layer, an activation layer, etc.
[0151] In the embodiments of the present disclosure, by jointly inputting the administrative level label corresponding to the candidate address word, the label anomaly probability corresponding to the candidate address word, and the word encoding vector into the semantic deviation recognition module, the semantic deviation module can combine the multi-dimensional features of the candidate address word to predict the semantic deviation type of the candidate address word, further improving the prediction effect of the semantic deviation type.
[0152] For ease of understanding, the following takes an example to describe the technical implementation details in the above embodiments:
[0153] In the telemarketing scenario, enterprise outbound robots usually need to obtain or confirm the address of the user. Due to the oral expression problems of the user and the escape problems of automatic speech recognition (ASR), the accuracy of address recognition will be affected. Moreover, when recognizing addresses, it is usually based on a single model for recognition. The ability of a single model is limited. Small models have insufficient ability to understand the context, while large models lack a complete address library and cannot achieve fine-grained recognition of addresses. Specifically, when dealing with complex contexts (oral expression problems, ASR escape problems) and multi-round long conversations, small models may not be able to accurately understand the context information, resulting in incorrect address recognition. Although large models have strong language processing capabilities, without the support of a dedicated knowledge base related to addresses, they cannot accurately identify address information, especially the corresponding relationship between address names and the levels of provinces, cities, districts, and counties. In addition, the existing address recognition methods cannot verify the recognition results, and directly returning the address recognition results is likely to affect the business quality, resulting in poor credibility of the address recognition results.
[0154] In view of this, the embodiments of the present disclosure provide an address recognition method, an address recognition device, an electronic device, and a storage medium, which can perform fused address recognition based on two models with different characteristics, effectively improving the accuracy of address recognition.
[0155] Figure 4 shows a schematic diagram of an address recognition method provided by an embodiment of the present disclosure. Referring to Figure 4 , the method includes:
[0156] Step S401: Read the conversation text.
[0157] Among them, the dialogue text can be one or more rounds of dialogue texts generated by one or more rounds of interactions between different call objects. For example:
[0158] Read the dialogue text text between the robot seat and the user in the system, that is, the ASR transcription result, which can be read once at the end of each round of dialogue. An example of the text text_i of the i-th round of dialogue (such as a question and answer) is as follows:
[0159] text_i (i = 1) = "Seat: Hello, I am xx Decoration Company. May I ask if you are Ms. Li? User: Yes.";
[0160] text_i (i = 2) = "Seat: Which district or county in Beijing is your house in? User: No, it's in Tianjing.";
[0161] text_i (i = 3) = "Seat: Where did you say? User: Binhai Binhai New Area.";
[0162] Among them, in the example of text_2, "Tianjing" is a semantic deviation caused by the customer's accent or habit problem. In the example of text_3, "Binhai Binhai New Area" is a semantic deviation due to the short pause in the user's speech and the lack of a comma after ASR transcription.
[0163] Step S402: Address-related judgment.
[0164] In the embodiments of the present disclosure, the second model can be used to perform address-related judgment on the dialogue text to extract the target dialogue text. Different from the traditional address-related rule judgment by keyword matching, the embodiments of the present disclosure use the second model (such as Qwen, GPT, etc.) to classify the dialogue text, so as to judge whether the dialogue text contains address-related information (especially when abbreviations with omitted words such as "province", "city", "district / county", etc. are involved), and output yes / no to improve the response efficiency of the second model.
[0165] For example, the prompt information prompt_assess_i for performing address-related judgment on the i-th round of dialogue text in the latest Q&A is as follows:
[0166] prompt_assess_i (i = 3) =
[0167] "Task: Please judge whether the current dialogue is related to the address according to the dialogue content.
[0168] Dialogue: Seat: Where did you say? User: Binhai Binhai New Area.
[0169] Output judgment result: yes or no (yes means related, no means not related)".
[0170] By inputting the prompt information into the second model, the output result output_i of the second model is obtained. For example, output_i (i = 3) is "Yes".
[0171] If the second model determines that the dialogue text has nothing to do with the address, step S406 can be executed to transfer to other processes for processing. If the second model determines that the dialogue text is related to the address, the target dialogue text can be generated according to the dialogue text, and step S403 can be executed to extract the address information.
[0172] Step S403: Perform address recognition according to the first model to obtain the address information.
[0173] Specifically, the target dialogue text can be input into the first model to extract the target address information corresponding to the target dialogue text according to the first model.
[0174] Among them, the first model is the traditional NLP model, which is a small model that does not perform pre-training processing. The first model can recognize the address information in the dialogue text of the user's answer in the i-th round of the latest Q&A, so as to extract information such as province, city, district, and county, and obtain the target address information.
[0175] The first model is a model specially trained for address recognition, which focuses on the extraction of address information. Exemplarily, the recognition result address_i of the first model for the i-th round of dialogue text text_i is as follows:
[0176] address_i (i = 2) = [{"name": "Address", "value": "Tianjing"}];
[0177] address_i (i = 3) = [{"name": "Address", "value": "Binhai Binhai New Area"}].
[0178] In an optional implementation manner, the first model can be an entity recognition model trained according to the dialogue text including address categories and non-address categories. The first model can mark the words belonging to the address category in the dialogue text. For example, "Tianjin" in the input "No, in Tianjin." will be marked as the address loc label.
[0179] Furthermore, the first model can combine the built-in preset address word library and post-processing rules to generate structured target address information according to the extracted address text words. For example, {"Province": "XX Province", "City": "XX City"}.
[0180] The preset address word library contains the administrative region types corresponding to the preset address words and the hierarchical relationships to which the preset address words belong, such as the corresponding hierarchical relationships between the entity names of provinces, cities, districts, and counties. The post-processing rules may include:
[0181] a. Determine whether the address text word is an entity corresponding to the provincial administrative region type. If so, output the entity name; otherwise, proceed to the next judgment.
[0182] b. Determine whether the address text word is an entity corresponding to the municipal administrative region type. If so, output the entity name; otherwise, proceed to the next judgment.
[0183] c. Determine whether the address text word is an entity corresponding to the district / county-level administrative region type. If so, output the entity name; otherwise, return an empty string.
[0184] d. If an entity corresponding to the district / county-level administrative region type is recognized, further determine whether entities corresponding to the municipal and provincial administrative region types are extracted. If not, supplement the entity names of the provincial and municipal administrative region types in combination with the preset address thesaurus. If so, directly return the entity names of the provincial, municipal, and district / county-level administrative region types.
[0185] Step S404: Verify the address information according to the second model, and correct the dialogue text according to the verification result.
[0186] Among them, the second model is determined according to the large language model. Since the large language model is an original open-source model, it is trained using the transformers architecture and self-supervised learning methods with a large amount of data. By masking part of the data, the model is prompted to learn the internal structure and features of the data, achieving efficient understanding and processing of the context. Therefore, in this step, the second model can correct semantic deviation problems such as ASR escape errors (i.e., 'Tianjin' is recognized as 'Tianjing' or others), and integrate the context, processing the discrete information of the multi-round dialogue text into a single-round dialogue text for the first model to perform structured address information extraction.
[0187] Exemplarily, the target address information address_i recognized by the first model (the results of i = 2, 3 are integrated) and the dialogue text text_i (the content of i = 2, 3 is integrated) are input to the second model together, and the second model (which can be an original large model without further training) performs verification and correction.
[0188] The prompt information prompt_verify_i for verifying and correcting the i-th round of dialogue text is as follows:
[0189] prompt_verify_i =
[0190] "Task: Please verify or correct the address recognition result according to the dialogue content.
[0191] Dialogue: Agent: Which district or county in Beijing is your house located in? User: No, it's in Tianjing. Agent: Where did you say? User: Binhai, Binhai New Area.
[0192] Address recognition result: [{"name": "Address", "value": "Tianjing"}, {"name": "Address", "value": "Binhai Binhai New Area"}].
[0193] Requirements:
[0194] 1) Output the discrimination result in json structure, with json primary keys being if_correct and corrected_text;
[0195] 2) If the recognition is correct, the value of if_correct is "Yes", and the value of corrected_text is an empty string;
[0196] 3) If the recognition is incorrect, the value of if_correct is "No", and the value of corrected_text is the corrected dialogue text.
[0197] By inputting the above prompt information into the second model, the expected output result can be obtained. For example, the output result is: "if_correct": "No", "corrected_text": "Agent: Which district or county in Beijing is your house located in? User: No, it's in Binhai, Tianjin's Binhai New Area."
[0198] If the verification result of the second model fails the verification, that is, the value of if_correct is "No" and the corrected dialogue text corrected_text is a non-empty value, then execute step S405.
[0199] Correspondingly, if the verification result of the second model passes the verification, that is, the value of if_correct is "Yes", then execute step S407.
[0200] Step S405: Correction times statistics.
[0201] In this step, the correction processing times of the second model can be recorded and statistically analyzed. For example, the count is k, and k is initialized to 0. Each time the verification result fails the verification and after the correction processing, the value of k is incremented by 1. If the number of corrections exceeds the preset threshold, then execute step S406 to transfer to other processings. If the address information extracted by the first model passes the verification of the second model after multiple address recognitions, then execute step S407.
[0202] Step S406: Transfer to other processes.
[0203] In this step, manual intervention can be used to obtain the corrected address information based on the target dialogue text, or other strategies can be adopted, such as using the address information corresponding to the target dialogue text after the last correction as the final corrected address information. This example does not limit this. For the dialogue text determined by the second model to be unrelated to the address, the dialogue text can also be processed accordingly according to other preset strategies.
[0204] Step S407: Return the address information.
[0205] Exemplarily, through several rounds of correction by the second model and address extraction by the first model, the final address information extraction result output is obtained. For example:
[0206] output = {
[0207] "Province": "A certain province",
[0208] "City": "A certain city",
[0209] "District": "A certain district"
[0210] }.
[0211] Thus, the system can return the obtained structured address recognition result, such as returning the extracted address recognition result to the user for address confirmation.
[0212] The address recognition method provided in this example has a certain corrective effect on semantic deviation noises such as accent problems and ASR escape problems in the dialogue text, thereby reducing address recognition errors caused by ASR escape and improving the robustness of address recognition. This example uses a strategy of combining multiple models, making full use of the language understanding ability of the large model and the fine recognition ability of the small model. The large model can understand the spoken language expression of the context, while the small model can accurately match the information in the address library for fine-grained address recognition. The combination of the two not only makes up for the deficiencies of a single model, but also significantly improves the overall accuracy of address recognition through mechanisms such as intelligent classification, multi-level recognition and verification, which helps to improve the business quality and user satisfaction.
[0213] It can be understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0214] In addition, the present disclosure also provides an address recognition device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any address recognition method provided by the present disclosure. For the corresponding technical solutions and descriptions, please refer to the corresponding records in the method section and will not be elaborated here.
[0215] Figure 5 It is a block diagram of an address recognition device provided by an embodiment of the present disclosure.
[0216] Referring to Figure 5 , an embodiment of the present disclosure provides an address recognition device, which includes:
[0217] An extraction module 51, configured to input a target dialogue text into a first model, and extract target address information corresponding to the target dialogue text according to the first model, where the first model is a trained address recognition model;
[0218] A verification module 52, configured to input the target address information and the target dialogue text into a second model, and verify the target address information according to the second model to obtain a verification result of the target address information, where the second model is determined according to a large language model;
[0219] A correction module 53, configured to, when the verification result fails the verification, perform correction processing on the target dialogue text according to the second model, so that the first model extracts corrected address information according to the corrected target dialogue text; and generate an address recognition result corresponding to the target dialogue text according to the corrected address information.
[0220] In an optional implementation manner, the performing correction processing on the target dialogue text according to the second model, so that the first model extracts corrected address information according to the corrected target dialogue text includes:
[0221] When the verification result of the i-th address information corresponding to the target dialogue text after the i-th correction does not meet the preset condition, perform the (i + 1)-th correction processing on the target dialogue text according to the second model to obtain the target dialogue text after the (i + 1)-th correction, where i is a natural number;
[0222] Input the target dialogue text after the (i + 1)-th correction into the first model, and extract the (i + 1)-th address information corresponding to the target dialogue text after the (i + 1)-th correction according to the first model;
[0223] Input the (i + 1)-th address information corresponding to the target dialogue text after the (i + 1)-th correction and the target dialogue text after the (i + 1)-th correction into the second model to obtain the verification result of the (i + 1)-th address information corresponding to the target dialogue text after the (i + 1)-th correction;
[0224] When the verification result of the (i + 1)-th address information corresponding to the target dialogue text after the (i + 1)-th correction meets the preset condition, obtain the corrected address information according to the (i + 1)-th address information corresponding to the target dialogue text after the (i + 1)-th correction; wherein, the preset condition includes: the verification result is verification passed and / or the number of correction processing times reaches the preset number threshold.
[0225] In an optional implementation manner, the (i + 1)-th correction processing of the target dialogue text by the second model includes:
[0226] Input a preset correction strategy into the second model, so that the second model performs the (i + 1)-th correction processing on the target dialogue text according to the preset correction strategy, wherein the preset correction strategy includes at least one of a dialogue merging strategy, a word order adjustment strategy, and a semantic adjustment strategy;
[0227] When the preset correction strategy is a dialogue merging strategy, the second model performing the (i + 1)-th correction processing on the target dialogue text according to the preset correction strategy includes: when the target dialogue text includes multi-round dialogue texts with context relationships, integrating the multi-round dialogue texts into one-round dialogue text according to the second model;
[0228] When the preset correction strategy is a word order adjustment strategy, the second model performing the (i + 1)-th correction processing on the target dialogue text according to the preset correction strategy includes: arranging the address information included in the target dialogue text in the order of the preset administrative region level according to the second model;
[0229] When the preset correction strategy is a semantic adjustment strategy, the second model performing the (i + 1)-th correction processing on the target dialogue text according to the preset correction strategy includes: performing semantic adjustment on the address information with semantic deviation included in the target dialogue text according to the second model.
[0230] In an optional implementation manner, the target dialogue text is obtained through the following method:
[0231] For any one-round dialogue text in the multi-round historical dialogue texts, identify whether the one-round dialogue text is a target dialogue text including address information according to the second model to obtain M initial target dialogue texts, where M is a natural number;
[0232] Determine whether there is a context relationship between the M initial target dialogue texts according to the second model, and splice at least two initial target dialogue texts with a context relationship to obtain N processed target dialogue texts, where N is a natural number less than or equal to M.
[0233] In an optional implementation manner, the extracting the target address information corresponding to the target dialogue text according to the first model includes:
[0234] Extract the address text words corresponding to the target dialogue text according to the first model;
[0235] Match the address text words with multiple preset address words included in a preset address word library, and determine the administrative region type corresponding to the address text words according to the matching result, where the preset address word library includes the administrative region type corresponding to the preset address words and the hierarchical relationship to which the preset address words belong;
[0236] When the administrative region type corresponding to the address text words is a preset type, determine the hierarchical relationship to which the address text words belong and the filled address words corresponding to the hierarchical relationship according to the hierarchical relationship to which the preset address words belong; perform a filling process on the address text words according to the filled address words to obtain the target address information.
[0237] In an optional implementation manner, the first model is trained according to the following method:
[0238] Perform semantic deviation processing on the standard address words included in the preset dialogue samples according to the preset semantic deviation types, to obtain the deviation dialogue samples corresponding to the preset dialogue samples, where the preset semantic deviation types at least include semantic repetition type, semantic pause type, and / or dialect conversion type;
[0239] Construct training samples according to the preset dialogue samples and the deviation dialogue samples, and input the training samples into an initial address recognition model to obtain the predicted address words corresponding to the training samples and the predicted semantic deviation types corresponding to the training samples;
[0240] Compare the standard address words corresponding to the training samples with the predicted address words corresponding to the training samples to obtain a first comparison result, and compare the preset semantic deviation types corresponding to the training samples with the predicted semantic deviation types to obtain a second comparison result;
[0241] Adjust the training parameters of the initial address recognition model according to the first comparison result and the second comparison result to obtain the first model.
[0242] In an alternative implementation, the apparatus is further configured to:
[0243] When the verification result is verification passed, generate an address recognition result corresponding to the target dialogue text according to the target address information.
[0244] Each module in the above address recognition apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0245] Figure 6 It is a block diagram of an electronic device provided by an embodiment of the present disclosure.
[0246] Referring to Figure 6 , an embodiment of the present disclosure provides an electronic device, which includes: at least one processor 601; at least one memory 602, and one or more I / O interfaces 603 connected between the processor 601 and the memory 602; wherein, the memory 602 stores one or more computer programs executable by at least one processor 601, and the one or more computer programs are executed by at least one processor 601 so that at least one processor 601 can execute the above address recognition method.
[0247] Each module in the above electronic device can be implemented in whole or in part by software, hardware, and their combination. The above respective modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0248] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, wherein the computer program implements the above address recognition method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0249] An embodiment of the present disclosure further provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of the electronic device, the processor in the electronic device executes the above address recognition method.
[0250] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components working together. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media).
[0251] As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically contains computer-readable program instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0252] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or can be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage media in each computing / processing device.
[0253] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the status information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0254] The computer program product described herein may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0255] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0256] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture, including instructions which implement various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0257] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0258] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special-purpose hardware and computer instructions.
[0259] Example embodiments have been disclosed herein, and although specific terms are employed, they are used in a general and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, the features, characteristics, and / or elements described in connection with a particular embodiment may be used singly or in combination with other embodiments. Accordingly, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. An address recognition method, characterized in that: include: Inputting the target conversation text into the first model, and extracting the target address information corresponding to the target conversation text according to the first model, wherein the first model is a trained address recognition model; Inputting the target address information and the target dialogue text into a second model, and verifying the target address information according to the second model to obtain a verification result of the target address information, wherein the second model is determined according to a large language model; When the verification result is failure to pass, the target dialogue text is corrected according to the second model so that the first model extracts corrected address information according to the corrected target dialogue text; and an address recognition result corresponding to the target dialogue text is generated according to the corrected address information.
2. The method according to claim 1, characterized in that The step of correcting the target dialogue text according to the second model so that the first model extracts the corrected address information according to the corrected target dialogue text includes: When the verification result of the i-th address information corresponding to the target dialogue text after the i-th correction does not meet the preset condition, the target dialogue text is corrected for the i+1th time according to the second model to obtain the target dialogue text after the i+1th correction, where i is a natural number; Inputting the target dialogue text after the (i+1)th correction into the first model, and extracting the (i+1)th address information corresponding to the target dialogue text after the (i+1)th correction according to the first model; Inputting the i+1th address information corresponding to the target dialogue text after the i+1th correction and the target dialogue text after the i+1th correction into the second model, and obtaining a verification result of the i+1th address information corresponding to the target dialogue text after the i+1th correction; When the verification result of the i+1th address information corresponding to the target conversation text after the i+1th correction satisfies the preset condition, the corrected address information is obtained according to the i+1th address information corresponding to the target conversation text after the i+1th correction; wherein the preset condition includes: the verification result is that the verification is passed and / or the number of correction processing times reaches a preset number threshold.
3. The method according to claim 2, characterized in that The performing the (i+1)th correction process on the target dialogue text according to the second model comprises: Inputting a preset correction strategy into the second model, so that the second model performs an (i+1)th correction process on the target dialogue text according to the preset correction strategy, wherein the preset correction strategy includes at least one of a dialogue merging strategy, a word order adjustment strategy, and a semantic adjustment strategy; In the case where the preset correction strategy is a dialogue merging strategy, the second model performs an i+1th correction process on the target dialogue text according to the preset correction strategy, including: in the case where the target dialogue text includes multiple rounds of dialogue texts having a contextual relationship, integrating the multiple rounds of dialogue texts into one round of dialogue text according to the second model; In the case where the preset correction strategy is a word order adjustment strategy, the second model performs an i+1th correction process on the target dialogue text according to the preset correction strategy, including: arranging the address information contained in the target dialogue text according to a preset administrative district level order according to the second model; In the case where the preset correction strategy is a semantic adjustment strategy, the second model performs the i+1th correction processing on the target dialogue text according to the preset correction strategy, including: according to the second model, semantically adjusting the address information with semantic deviation contained in the target dialogue text.
4. The method according to any one of claims 1 to 3, characterized in that: The target dialogue text is obtained by: For any round of dialogue texts among the multiple rounds of historical dialogue texts, identifying whether the any round of dialogue texts is a target dialogue text containing address information according to the second model, and obtaining M initial target dialogue texts, where M is a natural number; Determine whether the M initial target dialogue texts have a contextual relationship according to the second model, and concatenate at least two initial target dialogue texts having a contextual relationship to obtain N processed target dialogue texts, where N is a natural number less than or equal to M.
5. The method according to any one of claims 1 to 3, characterized in that: The step of extracting target address information corresponding to the target conversation text according to the first model includes: Extracting address text words corresponding to the target conversation text according to the first model; Matching the address text word with a plurality of preset address words contained in a preset address word library, and determining the administrative district type corresponding to the address text word according to the matching result, wherein the preset address word library contains the administrative district type corresponding to the preset address word and the hierarchical relationship to which the preset address word belongs; When the administrative district type corresponding to the address text word is a preset type, the hierarchical relationship to which the address text word belongs and the fill-in address word corresponding to the hierarchical relationship are determined according to the hierarchical relationship to which the preset address word belongs; the address text word is filled in according to the fill-in address word to obtain the target address information.
6. The method according to any one of claims 1 to 3, characterized in that: The first model is trained according to the following method: Performing semantic deviation processing on the standard address words included in the preset dialogue sample according to the preset semantic deviation type to obtain a deviation dialogue sample corresponding to the preset dialogue sample, wherein the preset semantic deviation type at least includes a semantic repetition type, a semantic pause type and / or a dialect escape type; Constructing a training sample according to the preset conversation sample and the deviation conversation sample, inputting the training sample into an initial address recognition model, and obtaining a predicted address word corresponding to the training sample and a predicted semantic deviation type corresponding to the training sample; Comparing the standard address word corresponding to the training sample with the predicted address word corresponding to the training sample to obtain a first comparison result, and comparing the preset semantic deviation type corresponding to the training sample with the predicted semantic deviation type to obtain a second comparison result; The training parameters of the initial address recognition model are adjusted according to the first comparison result and the second comparison result to obtain the first model.
7. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: When the verification result is that the verification is passed, an address recognition result corresponding to the target conversation text is generated according to the target address information.
8. An address recognition device, characterized in that: include: An extraction module, used for inputting a target conversation text into a first model, and extracting target address information corresponding to the target conversation text according to the first model, wherein the first model is a trained address recognition model; a verification module, configured to input the target address information and the target dialogue text into a second model, verify the target address information according to the second model, and obtain a verification result of the target address information, wherein the second model is determined according to a large language model; A correction module is used to correct the target dialogue text according to the second model when the verification result is verification failure, so that the first model can extract corrected address information based on the corrected target dialogue text; and generate an address recognition result corresponding to the target dialogue text based on the corrected address information.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the address identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the address recognition method according to any one of claims 1 to 7.