An object recognition method, apparatus, device, storage medium, and program product

By combining image and voice data to identify invoice fields and using confidence fusion technology to generate the final field values, the problem of low accuracy in invoice field recognition is solved, and high accuracy recognition is achieved under different conditions.

CN122290571APending Publication Date: 2026-06-26CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in recognizing document fields, especially when the image is unclear or the field layout and font are different, which can easily lead to recognition errors.

Method used

By combining image and speech data for recognition, the image recognition model and speech recognition model are used to generate the first and second field values ​​and their confidence scores, respectively, and the final third field value is generated based on the confidence scores. The image and speech recognition results are then fused to improve accuracy.

Benefits of technology

It improves the accuracy of document field recognition, and through multimodal fusion technology, it can still accurately identify field content even when the image is unclear or the field layout is different.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290571A_ABST
    Figure CN122290571A_ABST
Patent Text Reader

Abstract

This invention provides an object recognition method, apparatus, device, storage medium, and program product, relating to the field of data processing technology. The method includes: acquiring image data and first voice data of an object to be recognized, wherein the first voice data describes a target field in the object to be recognized; recognizing a first field value of the target field based on the image data, and a first confidence level corresponding to the first field value; recognizing a second field value of the target field based on the first voice data, and a second confidence level corresponding to the second field value; and generating a third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level. This invention can improve the accuracy of field recognition on invoices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to an object recognition method, apparatus, device, storage medium, and program product. Background Technology

[0002] Recognizing the content of negotiable instruments is a key technology in the financial field. In related technologies, models are typically used to identify the content of different fields within an instrument's image. However, this recognition relies heavily on the quality of the captured image; unclear images and variations in field layout and font across different instruments can lead to errors in image recognition, resulting in low accuracy.

[0003] It is evident that the relevant technologies suffer from low accuracy in recognizing fields on invoices. Summary of the Invention

[0004] This invention provides an object recognition method, apparatus, device, storage medium, and program product to solve the problem of low accuracy in field recognition of invoices in related technologies.

[0005] To solve the above problems, the present invention is implemented as follows:

[0006] In a first aspect, embodiments of the present invention provide an object recognition method, comprising:

[0007] Acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field situation in the object to be identified;

[0008] Based on the image data, identify the first field value of the target field and the first confidence level corresponding to the first field value;

[0009] Based on the first voice data, identify the second field value of the target field and the second confidence level corresponding to the second field value;

[0010] A third field value for the target field is generated based on the first field value, the first confidence level, the second field value, and the second confidence level. The third field value is the field value of the target field for the object to be identified.

[0011] Secondly, embodiments of the present invention also provide an object recognition device, comprising:

[0012] The acquisition module is used to acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field information in the object to be identified;

[0013] The first recognition module is used to recognize the first field value of the target field and the first confidence level corresponding to the first field value based on the image data.

[0014] The second recognition module is used to recognize the second field value of the target field and the second confidence level corresponding to the second field value based on the first voice data;

[0015] The first generation module is used to generate a third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level.

[0016] Thirdly, embodiments of the present invention also provide an electronic device, including a transceiver and a processor.

[0017] The transceiver is used to acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field situation in the object to be identified.

[0018] The processor is configured to identify a first field value of the target field and a first confidence level corresponding to the first field value based on the image data.

[0019] The processor is further configured to identify a second field value of the target field and a second confidence level corresponding to the second field value based on the first voice data;

[0020] The processor is further configured to generate a third field value for the target field based on the first field value, the first confidence level, the second field value, and the second confidence level.

[0021] Fourthly, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the object recognition method described in the first aspect.

[0022] Fifthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the object recognition method described in the first aspect.

[0023] In a sixth aspect, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the object recognition method described in the first aspect.

[0024] In this embodiment of the invention, image data and first voice data of an object to be identified are acquired. The first voice data describes the target field in the object to be identified. A first field value of the target field and a first confidence level corresponding to the first field value are identified based on the image data. A second field value of the target field and a second confidence level corresponding to the second field value are identified based on the first voice data. A third field value of the target field is generated based on the first field value, the first confidence level, the second field value, and the second confidence level. Thus, by obtaining the first field value and the first confidence level by recognizing the image data, obtaining the second field value and the second confidence level by recognizing the first voice data, and generating the third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level, the field value is corrected using the first voice data. This makes the final third field value more accurate than the field values ​​obtained by directly recognizing the image in related technologies, effectively improving the accuracy of field recognition. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart of an object recognition method provided in an embodiment of the present invention;

[0027] Figure 2 This is one of the overall schematic diagrams of the object recognition and model optimization process provided in the embodiments of the present invention;

[0028] Figure 3 This is one of the flowcharts illustrating object recognition provided in this embodiment of the invention;

[0029] Figure 4 This is the second schematic diagram of the object recognition process provided in the embodiment of the present invention;

[0030] Figure 5 This is one of the schematic diagrams of the model optimization process provided in the embodiments of the present invention;

[0031] Figure 6 This is the second schematic diagram of the model optimization process provided in the embodiment of the present invention;

[0032] Figure 7 This is the second overall schematic diagram of the object recognition and model optimization process provided in this embodiment of the invention;

[0033] Figure 8This is a structural diagram of an object recognition device provided in an embodiment of the present invention;

[0034] Figure 9 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] Please see Figure 1 , Figure 1 This is a flowchart of an object recognition method provided by an embodiment of the present invention, such as... Figure 1 As shown, it includes the following steps:

[0037] Step 101: Obtain image data and first voice data of the object to be identified. The first voice data is used to describe the target field information of the object to be identified.

[0038] The objects to be identified are those whose fields need to be identified. These can be tickets or other objects that include fields, such as letters, papers, or articles.

[0039] The aforementioned image data consists of images captured of the object to be identified. The image data includes images of all fields within the object to be identified. By identifying the image data, the field values ​​corresponding to the fields in the image can be determined. The device used to capture the image of the object to be identified can be a user terminal or a data acquisition device including a capturing component, enabling batch image data acquisition of different objects to be identified.

[0040] The aforementioned first voice data is data provided by the user describing the target field in the object to be identified. The value of the target field can be determined through this first voice data. The device acquiring the first voice data can be a user terminal or a data acquisition device including a sound sensor, to enable the acquisition of different types of voice data.

[0041] It should be noted that the first voice data can be voice data collected synchronously when acquiring image data, or voice data collected during the recognition of image data. The first voice data can be used to supplement or update the field value of the target field.

[0042] In some implementations, the image data of the object to be identified includes multiple fields, and the target field can be one of these fields. For a single field, if only the field value for image data recognition exists, but no corresponding voice data exists, the field value for image data recognition can be directly used as the final field value. Alternatively, a first field value can be obtained by recognizing the image data, and a second field value can be obtained through the first field data recognition; ultimately, the second field value can be directly used as the final field value.

[0043] For example, such as Figure 2 and Figure 3 As shown, after the user uploads image data of the object to be recognized, the data acquisition mode can be determined according to the specific situation. If the user only uploads image data and does not input the first voice data, the image data is recognized only by Optical Character Recognition (OCR) technology. However, if the user inputs the first voice data (such as one round of speech or at least two rounds of speech), in addition to recognizing the image data by OCR technology, the first voice data also needs to be recognized. Then, the different recognition results (i.e., the first field value and the second field value) are fused in a multimodal manner to obtain the final third field value.

[0044] In some implementations, after receiving image data and first voice data, the type of ticket corresponding to the image data can be identified, and then an image recognition model for identification can be determined based on the ticket type.

[0045] Furthermore, after receiving image data, the ambient light can be detected through image brightness histogram analysis. If the image quality is insufficient, the user will be prompted via voice or interface to adjust the shooting angle or lighting to ensure the quality of image data acquisition.

[0046] Step 102: Identify the first field value of the target field and the first confidence level corresponding to the first field value based on the image data.

[0047] The first field value mentioned above is a field value obtained by recognizing image data, and the first confidence level mentioned above is the confidence level corresponding to the first field value obtained by recognition. The credibility of the first field value can be determined through the first confidence level.

[0048] It should be noted that the target field contains both the recognition result of the image data (i.e., the first field value) and the recognition result of the first speech data (the second field value mentioned below). In this case, it is necessary to introduce the confidence level corresponding to the recognition result so that the final field value can be determined based on the confidence level corresponding to different recognition results.

[0049] In some implementations, the first field value of the target field and the first confidence level corresponding to the first field value are identified based on the image data, which can be obtained by a pre-trained image recognition model (such as a Transformer architecture model).

[0050] Furthermore, before identifying the first field value of the target field based on the image data, the image data can be preprocessed. The preprocessing includes adjusting the format of the image data, reducing noise in the image data, and calibrating perspective distortion.

[0051] Step 103: Identify the second field value of the target field and the second confidence level corresponding to the second field value based on the first voice data.

[0052] The second field value is the field value obtained by recognizing the first voice data, and the second confidence level is the confidence level corresponding to the recognized second field value. The credibility of the second field value can be determined through the second confidence level.

[0053] In some implementations, the identification of the second field value of the target field based on the first speech data, and the second confidence level corresponding to the second field value, can be obtained by a pre-trained speech recognition model (such as an end-to-end (ASR) model).

[0054] It should be noted that after acquiring the image data, the user's voice input is continuously monitored to obtain the first voice data. The monitoring of the user's voice input can be represented by the following formula:

[0055] ;

[0056] In the formula, E(t) represents the input speech signal, s(t,n) represents the speech signal value at the nth sampling point in the t-th frame, and N represents the total number of sampling points. Where E(t) is greater than a set threshold... In this case, the collected speech is the first speech data. At this time, the first speech data is recognized to obtain the second field value and the second confidence level.

[0057] Step 104: Generate a third field value for the target field based on the first field value, the first confidence level, the second field value, and the second confidence level. The third field value is the field value of the target field of the object to be identified.

[0058] It should be noted that after obtaining the first field value, the first confidence level, the second field value, and the second confidence level, the third field value of the target field can be generated based on the first field value, the first confidence level, the second field value, and the second confidence level. The third field value is the final field value, that is, the third field value is the final recognition result of the target field.

[0059] Specifically, the value of the third field can be determined by directly comparing the first and second confidence levels. Specifically, if the first confidence level is greater than the second confidence level, the first field value is set as the third field value; if the first confidence level is less than the second confidence level, the second field value is set as the third field value.

[0060] Alternatively, the first and second confidence levels can be used as calculation coefficients, and the first and second field values ​​can be merged to obtain the third field value, as detailed in the following examples.

[0061] In this embodiment of the invention, image data and first voice data of an object to be identified are acquired. The first voice data describes the target field in the object to be identified. A first field value of the target field and a first confidence level corresponding to the first field value are identified based on the image data. A second field value of the target field and a second confidence level corresponding to the second field value are identified based on the first voice data. A third field value of the target field is generated based on the first field value, the first confidence level, the second field value, and the second confidence level. Thus, by obtaining the first field value and the first confidence level by recognizing the image data, obtaining the second field value and the second confidence level by recognizing the first voice data, and generating the third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level, the field value is corrected using the first voice data. This makes the final third field value more accurate than the field values ​​obtained by directly recognizing the image in related technologies, effectively improving the accuracy of field recognition.

[0062] In one embodiment, the first voice data includes at least two rounds of voice, and the step of identifying a second field value of the target field based on the first voice data, and a second confidence level corresponding to the second field value, includes:

[0063] Identify the initial recognition result and initial confidence level corresponding to the first target speech, wherein the first target speech is any round of speech other than the first round of speech in the at least two rounds of speech;

[0064] Based on the intermediate recognition result and intermediate confidence level corresponding to the second target speech, and the initial recognition result and initial confidence level corresponding to the first target speech, an intermediate recognition result of the first target speech is generated, and the second target speech is the previous round of speech of the first target speech;

[0065] The intermediate confidence level of the first target speech is calculated based on the intermediate confidence level of the second target speech and the initial confidence level of the first target speech.

[0066] Wherein, if the first target speech is the last round of the at least two rounds of speech, the intermediate recognition result of the first target speech is the second field value, and the intermediate confidence of the first target speech is the second confidence.

[0067] It should be noted that the first voice data consists of at least two rounds of speech, meaning the user describes the target field multiple times. Therefore, it is necessary to recognize each round of speech and determine the second field value and the second confidence level based on the recognition results of each round. Taking the first target speech as an example, the first target speech can be any round of speech from at least two rounds; that is, for each round of speech, the intermediate recognition result and intermediate confidence level corresponding to that round can be obtained using the method described for the first target speech.

[0068] Specifically, for the first target speech, the initial recognition result and initial confidence level corresponding to the first target speech are first obtained. Then, the intermediate recognition result and intermediate confidence level of the second target speech are used to correct the initial recognition result and initial confidence level of the first target speech, resulting in the intermediate recognition result and intermediate confidence level of the first target speech. The intermediate recognition result and intermediate confidence level corresponding to each round of speech are calculated sequentially. In the case of calculating the last round of speech, the intermediate recognition result corresponding to the last round of speech is the second field value, and the intermediate confidence level corresponding to the last round of speech is the second confidence level.

[0069] In addition, for the first round of speech, since there is no previous round of speech, the initial recognition result corresponding to the first round of speech can be directly used as the intermediate recognition result, and the initial confidence level corresponding to the first round of speech can be used as the intermediate confidence level.

[0070] In this embodiment of the invention, an initial recognition result and an initial confidence level corresponding to a first target speech are identified. The first target speech is any round of speech other than the first round of speech in the at least two rounds of speech. Based on the intermediate recognition result and intermediate confidence level corresponding to a second target speech, and the initial recognition result and initial confidence level corresponding to the first target speech, an intermediate recognition result of the first target speech is generated. The second target speech is the previous round of speech of the first target speech. Based on the intermediate confidence level corresponding to the second target speech and the initial confidence level of the first target speech, the intermediate confidence level of the first target speech is calculated. In this way, by using the initial recognition result and initial confidence level corresponding to each round of speech, and the intermediate recognition result and intermediate confidence level corresponding to the previous round of speech, the intermediate recognition result and intermediate confidence level corresponding to each round of speech can be calculated, thereby obtaining the second field value and the second confidence level.

[0071] For example, the recognition result of the first round of user input can be used as the baseline field value and assigned a high confidence level. If the subsequent two rounds of voice error correction input produce numerical differences for the same field, the system uses the difference parsing module to dynamically adjust the value based on contextual semantic understanding and confidence level, avoiding incorrect overwriting of field values ​​due to a single misrecognition.

[0072] In some implementations, generating an intermediate recognition result for the first target speech based on the intermediate recognition result and intermediate confidence level corresponding to the second target speech, and the initial recognition result and initial confidence level corresponding to the first target speech, can be achieved by comparing the intermediate confidence level corresponding to the second target speech with the initial confidence level corresponding to the first target speech. If the intermediate confidence level corresponding to the second target speech is greater than the initial confidence level corresponding to the first target speech, the intermediate recognition result corresponding to the second target speech is set as the intermediate recognition result of the first target speech; otherwise, if the intermediate confidence level corresponding to the second target speech is less than the initial confidence level corresponding to the first target speech, the initial recognition result corresponding to the first target speech is set as the intermediate recognition result of the first target speech.

[0073] In other embodiments, the intermediate recognition result of the first target speech is generated based on the intermediate recognition result, intermediate confidence level, initial recognition result, and initial confidence level of the first target speech. Alternatively, the intermediate confidence level and the initial confidence level of the first target speech can be used as calculation coefficients to fuse the intermediate recognition result and the initial recognition result of the first target speech. This is the same as generating the third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level. A similar calculation formula can be set up to calculate the intermediate recognition result of the first target speech, referring to the calculation formula of the first field value, the first confidence level, the second field value, and the second confidence level in this invention.

[0074] In one embodiment, calculating the intermediate confidence of the first target speech based on the intermediate confidence of the second target speech and the initial confidence of the first target speech includes:

[0075] The quotient of the first product and the first sum is set as the intermediate confidence level of the first target speech. The first sum is the sum of the first product and the second product. The first product is the product of the intermediate confidence level of the second target speech and the initial confidence level of the first target speech. The second product is the product of the first difference and the second difference. The first difference is the difference between the first preset coefficient and the intermediate confidence level of the second target speech. The second difference is the difference between the second preset coefficient and the initial confidence level of the first target speech.

[0076] In this embodiment of the invention, the quotient of the first product and the first sum is set as the intermediate confidence level of the first target speech. The first sum is the sum of the first product and the second product. The first product is the product of the intermediate confidence level of the second target speech and the initial confidence level of the first target speech. The second product is the product of the first difference and the second difference. The first difference is the difference between the first preset coefficient and the intermediate confidence level of the second target speech, and the second difference is the difference between the second preset coefficient and the initial confidence level of the first target speech. Thus, the intermediate confidence level of the first target speech is calculated using the above method.

[0077] Specifically, the intermediate confidence score of the first target speech can be calculated using the following formula:

[0078] ;

[0079] In the formula, c t(x) represents the median confidence score of the t-th round of speech (i.e., the first target speech), and c t-1 (x) represents the median confidence score of the (t-1)th round of speech (i.e., the second target speech). The initial confidence level of the first target speech is 1, with the first preset coefficient and the second preset coefficient being the same.

[0080] It should be noted that the first preset coefficient and the second preset coefficient in the formula can be the same or different, and the values ​​of the first preset coefficient and the second preset coefficient can be configured according to specific needs.

[0081] In this way, the above formula can be used to update the field values ​​and weights based on the new round of voice input, and take into account the previously input voice, so that the final recognized second field value can meet the user's intention.

[0082] In one embodiment, generating the third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level includes:

[0083] The quotient of the second sum and the third sum is set as the third field value. The second sum is the sum of the third product and the fourth product. The third product is the product of the first field value and the first confidence level. The fourth product is the product of the second field value and the second confidence level. The third sum is the sum of the first confidence level and the second confidence level.

[0084] In this embodiment of the invention, the quotient of the second and third sums is set as the third field value. The second sum is the sum of the third and fourth products. The third product is the product of the first field value and the first confidence level, and the fourth product is the product of the second field value and the second confidence level. The third sum is the sum of the first and second confidence levels. This achieves the fusion of the first and second field values ​​to obtain the third field value, which is more accurate than the first field value, which is only used for image data recognition.

[0085] Specifically, the value of the third field of the target field can be represented by the following formula:

[0086] ;

[0087] In the formula x fuse The value of the third field, x ocr c is the value of the first field. ocr For the first confidence level, x asr For the value of the second field, c asr This represents the second confidence level.

[0088] In one embodiment, the method further includes:

[0089] Display the value of the third field of the target field;

[0090] The first confidence level and the second confidence level are weighted to obtain the third confidence level;

[0091] If the third confidence level is less than a preset confidence threshold, the value of the third field will be highlighted.

[0092] It should be noted that after recognizing the third field value, it is also necessary to calculate the corresponding third confidence level to determine the reliability of the third field value. In some scenarios, even if the user inputs the first voice data, there may be cases where the third confidence level is low, such as when the first voice data is unclear or the image data is unclear. In such cases, even if the third field value is recognized, the third field value may not be accurate enough, and the user needs to make a judgment.

[0093] Therefore, in this embodiment of the invention, the third field value of the target field is displayed; the first confidence level and the second confidence level are weighted to obtain a third confidence level; when the third confidence level is less than a preset confidence threshold, the third field value is highlighted. In this way, by setting a preset confidence threshold, and highlighting the third field value when the third confidence level is less than the preset confidence threshold, users can quickly locate the highlighted third field value to determine whether the third field value needs to be adjusted.

[0094] In some implementations, the first confidence level and the second confidence level can be weighted to obtain a third confidence level. The weighting coefficients of the first and second confidence levels can be pre-configured, typically with the second confidence level having a larger weight to increase the influence of the first voice data on the third confidence level.

[0095] In one embodiment, after displaying the third field value of the target field, the method further includes:

[0096] Receive second voice data;

[0097] Identify the value of the fourth field corresponding to the second voice data;

[0098] Determine the field corresponding to the value of the fourth field;

[0099] If the field corresponding to the fourth field value is the target field, update the field value of the target field based on the fourth field value.

[0100] The aforementioned second voice data is the voice data received after the third field value of the target field is highlighted, and the second voice data is voice data specific to the target field. At this time, the target field can be updated based on the fourth field value obtained from the recognition of the second voice data to further improve the accuracy of target field recognition.

[0101] The identification of the fourth field value corresponding to the second speech data can be achieved using the same method as the identification of the first speech data. Specifically, the second speech data can be identified by a pre-trained speech recognition model to obtain the fourth field value.

[0102] Specifically, such as Figure 4 As shown, after recognizing the image data and the first voice data, the third field value is obtained, and then updated according to the second voice data to achieve multiple rounds of dynamic updates, so as to output the final field value.

[0103] It should be noted that after the fourth field value is identified, the field corresponding to the fourth field value can be determined based on the dynamic field priority algorithm and the updatable field mapping table. If the field corresponding to the fourth field value is the target field, the field value of the target field is updated based on the fourth field value.

[0104] The dynamic field priority algorithm can be expressed by the following formula:

[0105] ;

[0106] In the formula, k i f is the value of the i-th field (i.e., the value of the fourth field). j Let j be the j-th field, F be an updatable field mapping table, and match(k) = ... i ,f j ) represents the matching degree between the i-th field value and the j-th field; p j This indicates the priority of the j-th field, which can be adjusted based on at least two rounds of speech and historical data; The field that is finally matched for the value of the i-th field.

[0107] The updatable field mapping table can be represented as follows:

[0108]

[0109] In this way, a dynamic field priority algorithm and an updatable field mapping table are used to quickly determine the field corresponding to the fourth field value (i.e., the final mapping in the table).

[0110] The current priority is automatically adjusted by a dynamic field priority algorithm, taking into account the context and historical operations; the final mapping indicates that the system selects the field with the highest matching degree and the highest priority for performing ticket modification or supplement operations.

[0111] In this embodiment of the invention, second voice data is received; a fourth field value corresponding to the second voice data is identified; a field corresponding to the fourth field value is determined; and if the field corresponding to the fourth field value is the target field, the field value of the target field is updated based on the fourth field value. Thus, updating the target field based on the fourth field value obtained from identifying the second voice data further improves the accuracy of target field identification.

[0112] In some implementations, updating the target field value based on the fourth field value can be achieved by calculating the difference between the fourth field value and the third field value. If the difference is greater than a set difference threshold, the fourth field value is set as the target field value, or the third field value is supplemented based on the fourth field value. If the difference is less than or equal to the preset difference threshold, the third field value is still used as the target field value. This method, by determining whether to update the target field value using the difference value and the preset difference threshold, avoids frequent updates to the field value when the difference is small, thus ensuring a relatively stable final field value.

[0113] For example, a difference value greater than a set difference threshold can be expressed as:

[0114] ;

[0115] in the formula For the difference value, v user The value of the fourth field, v sys This is the value of the third field.

[0116] Furthermore, after setting the fourth field value as the target field value, it is also necessary to adjust the confidence level of the target field to make it closer to the second voice data input by the user. The confidence level of the target field can be adjusted using the following formula:

[0117] ;

[0118] In the formula, c new c is the adjusted confidence level for the target field. old c represents the confidence level of the target field before adjustment. user The confidence level of the fourth field value is given by 'a', where 'a' is the weighting coefficient, and the value of 'a' ranges from 0 to 1.

[0119] Furthermore, if the difference value is greater than a set difference threshold, the value of the fourth field can be set as the value of the target field, or the value of the third field can be supplemented based on the value of the fourth field. This can be expressed by the following formula:

[0120] ;

[0121] In the formula f intent ( ) represents the semantic intent classification function, T is the value of the fourth field, and I is the update operation. That is, the value of the fourth field can be used to replace, delete, or supplement the value of the target field.

[0122] In some implementations, second voice data can also be acquired through continuous monitoring and low-power wake-up. Specifically, for example... Figure 5 As shown, continuous monitoring and low-power wake-up are used to collect second voice data. After the second voice data is collected, the fourth field value corresponding to the second voice data is identified, and the field corresponding to the fourth field value is determined and updated. This achieves the correction of the field value by at least two rounds of voice data. Subsequently, the recognition model is iteratively updated to optimize it. The optimized model is then used to further identify the data collected by the subsequent continuous monitoring and low-power wake-up methods to improve the recognition accuracy.

[0123] One method for low-power wake-up is to monitor and wake up the user by detecting keywords. Specifically, this involves identifying keywords indicating the user's intention to correct errors, such as "not," "change," or "modify." Low-power wake-up can be represented by the following formula:

[0124] ;

[0125] y in the formula L (t) represents the recognition result of the low-power keyword detection function, f L ( ) is a low-power keyword detection function, y L (t)=1 indicates that the keyword was detected, y L x(t) = 0 indicates that no keyword was detected, and x(t) represents the collected vocabulary.

[0126] Furthermore, after detecting keywords, a complete speech analysis is performed through a high-precision automatic speech recognition (ASR) channel, specifically expressed by the following formula:

[0127] ;

[0128] y in the formula H (t) represents the recognized text obtained by the high-precision speech recognition function, f H ( ) represents a high-precision speech recognition function, and T represents the recognized text y. H The set of (t).

[0129] That is, in this invention, if y L When (t)=1, switch to the high-precision speech recognition function f H ( ), otherwise keep the low-power keyword detection function f L ( ).

[0130] In one embodiment, when at least two of the second voice data are received, the time interval between the reception times of the at least two second voice data is greater than a preset interval threshold.

[0131] In this embodiment of the invention, to avoid situations where frequent input of voice data by users leads to insufficient recognition time or high recognition delay, a preset interval threshold is introduced. By using the preset interval threshold to avoid frequent input, the updated field value of the target field can better match the content of the user's input voice data.

[0132] It should be noted that after identifying the final field value of the target field (such as the third or fourth field value mentioned above), the final field value needs to be displayed for user confirmation. As mentioned throughout the text, in addition to avoiding frequent changes in field values ​​due to frequent voice data input through a preset interval threshold, the difference between the fourth and third field values ​​and a preset difference threshold can be used to determine whether the target field value needs to be updated. Specifically, the following formula can be used to determine whether the displayed target field value needs to be updated:

[0133] ;

[0134] In the formula, F(t) indicates whether an update is triggered. An update is triggered when F(t) = 1, and no update is performed when F(t) = 0; t is the current time of voice data reception. last The last time voice data was received. For the preset interval threshold, This represents the difference between the values ​​in the fourth and third fields. This is a preset difference threshold.

[0135] It should be noted that the set of fields of the object to be identified can be represented as follows: The number of rounds of voice data is The process requires updating the fields of the object to be recognized based on all received voice data in order to complete the field recognition. This can be represented by the following formula:

[0136] ;

[0137] If C rIf <1, it indicates there are still fields to be processed, and the system will continue to guide the user to speak in the next round of dialogue; if C r =1 indicates that all fields have been corrected, and the system ends the multi-round interaction.

[0138] Guiding user comments includes the following methods:

[0139] Voice prompts: For fields with low confidence or ambiguity, the system uses speech synthesis technology to proactively announce the recognition results and potential problems to the user, prompting them to confirm or make corrections. For example: "The invoice amount recognition result is uncertain. Please verify whether the amount is correct."

[0140] Field highlighting and interactive operations: The front-end highlights fields with low confidence or detected anomalies, allowing users to directly click to enter the editing interface for modification, improving correction efficiency. Simultaneously, the highlighted field displays related contextual information to help users fully understand the field's status.

[0141] Input suggestions and auto-completion: By combining contextual semantics and historical user input data, the system intelligently recommends possible input content or formats to help users quickly complete field filling and reduce the input error rate.

[0142] Multi-round error correction guidance: When users make voice error corrections or manual modifications, the system provides real-time feedback on the latest recognition status and dynamically adjusts the guidance strategy based on the correction results, achieving closed-loop optimization of multi-round interaction.

[0143] This dynamic guidance mechanism does not follow a fixed, pre-set process. Instead, it dynamically adjusts based on real-time calculated confidence levels and contextual semantics, exhibiting high flexibility and intelligence. It can adapt to different ticket types and complexities, meet diverse user operating habits, and significantly improve the overall recognition accuracy and user satisfaction.

[0144] Furthermore, such as Figure 6 As shown, after highlighting the field value, a preset interval threshold is set to avoid frequent changes, and the user is guided to input second voice data. The fourth field value is obtained through the recognition of the second voice data and the field value is updated. Furthermore, the third and fourth field values ​​can be collected to create training samples to train the recognition model.

[0145] Specifically, in one embodiment, the step of identifying the first field value of the target field based on the image data, and the first confidence level corresponding to the first field value, includes:

[0146] The image data is identified based on a preset image recognition model to obtain the first field value of the target field and the first confidence level corresponding to the first field value;

[0147] The step of identifying the second field value of the target field based on the first speech data, and the second confidence level corresponding to the second field value, includes:

[0148] The first speech data is recognized based on a preset speech recognition model to obtain the second field value of the target field and the second confidence level corresponding to the second field value;

[0149] The method further includes:

[0150] Sample data is generated based on the values ​​of the third and fourth fields of multiple fields, wherein the multiple fields include the target field;

[0151] The preset image recognition model is trained based on the sample data to obtain the optimized preset image recognition model;

[0152] The preset speech recognition model is trained based on the sample data to obtain the optimized preset speech recognition model.

[0153] The aforementioned preset image recognition model is a model for recognizing image data, and the aforementioned preset speech recognition model is a model for recognizing the first set of data. In this embodiment of the invention, sample data is generated based on the values ​​of the third and fourth fields of multiple fields, and then the preset image recognition model and the preset speech recognition model are optimized using the sample data to improve the recognition accuracy of the optimized preset image recognition model and the preset speech recognition model.

[0154] Specifically, the preset image recognition model is trained based on the sample data to obtain an optimized preset image recognition model, or the preset speech recognition model is trained based on the sample data to obtain an optimized preset speech recognition model, which can be expressed by the following formula:

[0155] ;

[0156] in the formula Let be the parameters of the model during the t-th training iteration (i.e., before optimization). Let R be the parameters of the model trained (i.e., after optimization) for the (t+1)th time. orig For the value of the third field, R corr The fourth field value is L(), where L() is the loss function. For learning rate, is a coefficient.

[0157] In this way, the above formula can be used to optimize the preset image recognition model and the preset speech recognition model.

[0158] Furthermore, the process of identifying fields and optimizing models is as follows: Figure 7 As shown. Figure 7 As shown, the recognition results are processed by field extraction and format standardization to obtain structured data (i.e., the third field value); the data is then displayed, and the corrected results (i.e., the fourth field value) are obtained. Differences are compared, and sample data is constructed to optimize the recognition model. The optimized model can then be used for further recognition to improve the recognition accuracy.

[0159] The differences can be represented by the following formula:

[0160] ;

[0161] Result in the formula final For the final identification result (i.e., the value of the fourth field), Result original This is the initial recognition result (i.e., the value of the third field). The difference value.

[0162] For example, when a user uploads a complex travel invoice, the system immediately begins image analysis after the photo is uploaded, simultaneously activating continuous monitoring. Regardless of when the user speaks, the system can instantly capture voice commands without requiring the user to manually switch voice modes or click confirm. The system first preprocesses and performs initial recognition on the invoice, extracting key fields (i.e., third-field values) such as invoice number, amount, date, and merchant name. The preliminary results are then returned to the front-end interface, highlighting low-confidence fields and providing voice prompts to alert the user to potentially requiring confirmation.

[0163] Due to the complexity of the information on the invoice, the system engages in multiple rounds of dialogue with the user. In the first round, the system indicates that the amount field may be inaccurate. After the user corrects the amount via voice (i.e., the value of the fourth field), the system immediately updates the corresponding field and continues to listen for subsequent voice commands. In the second round, the system prompts that the accommodation date and invoice number need to be confirmed. The user confirms or corrects them one by one via voice, and the system updates the recognition status in the background in real time. Even if the user discovers a spelling error in the merchant name during the dialogue, they can directly correct it via voice, and the system will process it immediately within the current session without interruption or re-uploading.

[0164] After all fields are confirmed, the system generates structured information from the final recognition result and records every correction action taken by the user. This feedback data is fed back into the backend learning module to optimize the recognition model and speech parsing capabilities, enabling the system to recognize similar documents more quickly and accurately in the future. In this way, the entire process achieves a complete closed loop, from user input, backend multimodal fusion, frontend real-time feedback, at least two rounds of speech error correction, to structured generation and closed-loop optimization. This not only ensures high recognition accuracy but also provides a natural and efficient user interaction experience.

[0165] Please see Figure 8 , Figure 8 This is a structural diagram of an object recognition device provided in an embodiment of the present invention, as shown below. Figure 8 As shown, the object recognition device 800 includes:

[0166] The acquisition module 801 is used to acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field situation in the object to be identified.

[0167] The first recognition module 802 is used to recognize the first field value of the target field and the first confidence level corresponding to the first field value based on the image data.

[0168] The second recognition module 803 is used to recognize the second field value of the target field and the second confidence level corresponding to the second field value based on the first voice data.

[0169] The first generation module 804 is used to generate a third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level, wherein the third field value is the field value of the target field of the object to be identified.

[0170] In one embodiment, the first voice data includes at least two rounds of voice, and the second recognition module 803 includes:

[0171] The first recognition submodule is used to recognize the initial recognition result and initial confidence level corresponding to the first target speech, wherein the first target speech is any round of speech other than the first round of speech in the at least two rounds of speech;

[0172] The generation submodule is used to generate the intermediate recognition result of the first target speech based on the intermediate recognition result and intermediate confidence of the second target speech, the initial recognition result and initial confidence of the first target speech, and the second target speech is the previous round of speech of the first target speech;

[0173] The calculation submodule is used to calculate the intermediate confidence of the first target speech based on the intermediate confidence of the second target speech and the initial confidence of the first target speech;

[0174] Wherein, if the first target speech is the last round of the at least two rounds of speech, the intermediate recognition result of the first target speech is the second field value, and the intermediate confidence of the first target speech is the second confidence.

[0175] In one embodiment, the computing submodule includes:

[0176] The first calculation unit is configured to set the quotient of the first product and the first sum as the intermediate confidence level of the first target speech. The first sum is the sum of the first product and the second product. The first product is the product of the intermediate confidence level of the second target speech and the initial confidence level of the first target speech. The second product is the product of the first difference and the second difference. The first difference is the difference between the first preset coefficient and the intermediate confidence level of the second target speech. The second difference is the difference between the second preset coefficient and the initial confidence level of the first target speech.

[0177] In one embodiment, the first generation module 804 includes:

[0178] The second calculation unit is used to set the quotient of the second sum and the third sum as the third field value, wherein the second sum is the sum of the third product and the fourth product, the third product is the product of the first field value and the first confidence level, the fourth product is the product of the second field value and the second confidence level, and the third sum is the sum of the first confidence level and the second confidence level.

[0179] In one embodiment, the object recognition device 800 further includes:

[0180] The display module is used to display the value of the third field of the target field;

[0181] The weighting module is used to weight the first confidence level and the second confidence level to obtain the third confidence level;

[0182] The highlighting module is used to highlight the value of the third field when the third confidence level is less than a preset confidence threshold.

[0183] In one embodiment, the object recognition device 800 further includes:

[0184] The receiving module is used to receive the second voice data;

[0185] The third recognition module is used to recognize the fourth field value corresponding to the second voice data;

[0186] The determination module is used to determine the field corresponding to the value of the fourth field;

[0187] An update module is used to update the field value of the target field based on the fourth field value when the field corresponding to the fourth field value is the target field.

[0188] In one embodiment, when at least two of the second voice data are received, the time interval between the reception times of the at least two second voice data is greater than a preset interval threshold.

[0189] In one embodiment, the first identification module 802 includes:

[0190] The second recognition submodule uses a preset image recognition model to recognize the image data and obtain the first field value of the target field and the first confidence level corresponding to the first field value.

[0191] The second identification module 803 includes:

[0192] The third recognition submodule recognizes the first speech data based on a preset speech recognition model to obtain the second field value of the target field and the second confidence level corresponding to the second field value.

[0193] The object recognition device 800 further includes:

[0194] The second generation module is used to generate sample data based on the values ​​of the third and fourth fields of multiple fields, wherein the multiple fields include the target field;

[0195] The first training module is used to train the preset image recognition model based on the sample data to obtain the optimized preset image recognition model.

[0196] The second training module is used to train the preset speech recognition model based on the sample data to obtain the optimized preset speech recognition model.

[0197] The object recognition device provided in this embodiment of the invention can implement each process of the above-described object recognition method, with one-to-one correspondence of technical features and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0198] It should be noted that the object recognition device in the embodiments of the present invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0199] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the above-described functionality. Figure 1 The various processes of the object recognition method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0200] For details, see Figure 9 As shown, this embodiment of the invention also provides an electronic device, including a bus 901, a transceiver 902, an antenna 903, a bus interface 904, a processor 905, and a memory 906.

[0201] The transceiver 902 is used to acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field situation in the object to be identified.

[0202] The processor 905 is configured to identify a first field value of the target field and a first confidence level corresponding to the first field value based on the image data.

[0203] The processor 905 is further configured to identify a second field value of the target field and a second confidence level corresponding to the second field value based on the first voice data;

[0204] The processor 905 is further configured to generate a third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level, wherein the third field value is the field value of the target field of the object to be identified.

[0205] In one embodiment, the first voice data includes at least two rounds of voice, and the step of identifying a second field value of the target field based on the first voice data, and a second confidence level corresponding to the second field value, includes:

[0206] Identify the initial recognition result and initial confidence level corresponding to the first target speech, wherein the first target speech is any round of speech other than the first round of speech in the at least two rounds of speech;

[0207] Based on the intermediate recognition result and intermediate confidence level corresponding to the second target speech, and the initial recognition result and initial confidence level corresponding to the first target speech, an intermediate recognition result of the first target speech is generated, and the second target speech is the previous round of speech of the first target speech;

[0208] The intermediate confidence level of the first target speech is calculated based on the intermediate confidence level of the second target speech and the initial confidence level of the first target speech.

[0209] Wherein, if the first target speech is the last round of the at least two rounds of speech, the intermediate recognition result of the first target speech is the second field value, and the intermediate confidence of the first target speech is the second confidence.

[0210] In one embodiment, calculating the intermediate confidence of the first target speech based on the intermediate confidence of the second target speech and the initial confidence of the first target speech includes:

[0211] The quotient of the first product and the first sum is set as the intermediate confidence level of the first target speech. The first sum is the sum of the first product and the second product. The first product is the product of the intermediate confidence level of the second target speech and the initial confidence level of the first target speech. The second product is the product of the first difference and the second difference. The first difference is the difference between the first preset coefficient and the intermediate confidence level of the second target speech. The second difference is the difference between the second preset coefficient and the initial confidence level of the first target speech.

[0212] In one embodiment, generating the third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level includes:

[0213] The quotient of the second sum and the third sum is set as the third field value. The second sum is the sum of the third product and the fourth product. The third product is the product of the first field value and the first confidence level. The fourth product is the product of the second field value and the second confidence level. The third sum is the sum of the first confidence level and the second confidence level.

[0214] In one embodiment, the processor 905 is further configured to display the value of a third field of the target field;

[0215] The processor 905 is further configured to weight the first confidence level and the second confidence level to obtain a third confidence level;

[0216] The processor 905 is further configured to highlight the third field value when the third confidence level is less than a preset confidence threshold.

[0217] In one embodiment, the transceiver 902 is further configured to receive second voice data;

[0218] The processor 905 is also used to identify the fourth field value corresponding to the second voice data;

[0219] The processor 905 is further configured to determine the field corresponding to the fourth field value;

[0220] The processor 905 is further configured to update the field value of the target field based on the fourth field value when the field corresponding to the fourth field value is the target field.

[0221] In one embodiment, when at least two of the second voice data are received, the time interval between the reception times of the at least two second voice data is greater than a preset interval threshold.

[0222] In one embodiment, identifying the first field value of the target field based on the image data, and the first confidence level corresponding to the first field value, includes:

[0223] The image data is identified based on a preset image recognition model to obtain the first field value of the target field and the first confidence level corresponding to the first field value;

[0224] The step of identifying the second field value of the target field based on the first speech data, and the second confidence level corresponding to the second field value, includes:

[0225] The first speech data is recognized based on a preset speech recognition model to obtain the second field value of the target field and the second confidence level corresponding to the second field value;

[0226] The processor 905 is further configured to generate sample data based on the values ​​of a third field and a fourth field of a plurality of fields, the plurality of fields including the target field;

[0227] The processor 905 is further configured to train the preset image recognition model based on the sample data to obtain an optimized preset image recognition model;

[0228] The processor 905 is further configured to train the preset speech recognition model based on the sample data to obtain an optimized preset speech recognition model.

[0229] exist Figure 9In this document, a bus architecture (represented by bus 901) is used. Bus 901 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 905 and memory represented by memory 906. Bus 901 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 904 provides an interface between bus 901 and transceiver 902. Transceiver 902 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 905 is transmitted over a wireless medium via antenna 903, which further receives data and transmits it to processor 905.

[0230] Processor 905 manages bus 901 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 906 can be used to store data used by processor 905 during operation.

[0231] Optionally, the processor 905 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a graphics processing unit (GPU).

[0232] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the above-described functions. Figure 1 The various processes of the corresponding object recognition method embodiments, which achieve the same technical effect, will not be described again here to avoid repetition. The computer-readable storage medium mentioned includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0233] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The various processes of the corresponding object recognition method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.

[0234] In the embodiments of this invention, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0235] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0236] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0237] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An object recognition method, characterized in that, include: Acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field situation in the object to be identified; Based on the image data, identify the first field value of the target field and the first confidence level corresponding to the first field value; Based on the first voice data, identify the second field value of the target field and the second confidence level corresponding to the second field value; A third field value for the target field is generated based on the first field value, the first confidence level, the second field value, and the second confidence level. The third field value is the field value of the target field for the object to be identified.

2. The method as described in claim 1, characterized in that, The first voice data includes at least two rounds of speech. The step of identifying the second field value of the target field based on the first voice data, and the second confidence level corresponding to the second field value, includes: Identify the initial recognition result and initial confidence level corresponding to the first target speech, wherein the first target speech is any round of speech other than the first round of speech in the at least two rounds of speech; Based on the intermediate recognition result and intermediate confidence level corresponding to the second target speech, and the initial recognition result and initial confidence level corresponding to the first target speech, an intermediate recognition result of the first target speech is generated, and the second target speech is the previous round of speech of the first target speech; The intermediate confidence level of the first target speech is calculated based on the intermediate confidence level of the second target speech and the initial confidence level of the first target speech. Wherein, if the first target speech is the last round of the at least two rounds of speech, the intermediate recognition result of the first target speech is the second field value, and the intermediate confidence of the first target speech is the second confidence.

3. The method as described in claim 2, characterized in that, The step of calculating the intermediate confidence level of the first target speech based on the intermediate confidence level corresponding to the second target speech and the initial confidence level of the first target speech includes: The quotient of the first product and the first sum is set as the intermediate confidence level of the first target speech. The first sum is the sum of the first product and the second product. The first product is the product of the intermediate confidence level of the second target speech and the initial confidence level of the first target speech. The second product is the product of the first difference and the second difference. The first difference is the difference between the first preset coefficient and the intermediate confidence level of the second target speech. The second difference is the difference between the second preset coefficient and the initial confidence level of the first target speech.

4. The method as described in claim 1, characterized in that, The step of generating the third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level includes: The quotient of the second sum and the third sum is set as the third field value. The second sum is the sum of the third product and the fourth product. The third product is the product of the first field value and the first confidence level. The fourth product is the product of the second field value and the second confidence level. The third sum is the sum of the first confidence level and the second confidence level.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Display the value of the third field of the target field; The first confidence level and the second confidence level are weighted to obtain the third confidence level; If the third confidence level is less than a preset confidence threshold, the value of the third field will be highlighted.

6. The method as described in claim 5, characterized in that, After displaying the third field value of the target field, the method further includes: Receive second voice data; Identify the value of the fourth field corresponding to the second voice data; Determine the field corresponding to the value of the fourth field; If the field corresponding to the fourth field value is the target field, update the field value of the target field based on the fourth field value.

7. The method as described in claim 6, characterized in that, When at least two of the second voice data are received, the time interval between the reception times of the at least two second voice data is greater than a preset interval threshold.

8. The method as described in claim 6, characterized in that, The step of identifying the first field value of the target field based on the image data, and the first confidence level corresponding to the first field value, includes: The image data is identified based on a preset image recognition model to obtain the first field value of the target field and the first confidence level corresponding to the first field value; The step of identifying the second field value of the target field based on the first speech data, and the second confidence level corresponding to the second field value, includes: The first speech data is recognized based on a preset speech recognition model to obtain the second field value of the target field and the second confidence level corresponding to the second field value; The method further includes: Sample data is generated based on the values ​​of the third and fourth fields of multiple fields, wherein the multiple fields include the target field; The preset image recognition model is trained based on the sample data to obtain the optimized preset image recognition model; The preset speech recognition model is trained based on the sample data to obtain the optimized preset speech recognition model.

9. An object recognition device, characterized in that, include: The acquisition module is used to acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field information in the object to be identified; The first recognition module is used to recognize the first field value of the target field and the first confidence level corresponding to the first field value based on the image data. The second recognition module is used to recognize the second field value of the target field and the second confidence level corresponding to the second field value based on the first voice data; The first generation module is used to generate a third field value of the target field based on the first field value, the first confidence level, the second field value, and the second confidence level, wherein the third field value is the field value of the target field of the object to be identified.

10. An electronic device, characterized in that, Including transceivers and processors, The transceiver is used to acquire image data and first voice data of the object to be identified, wherein the first voice data is used to describe the target field situation in the object to be identified. The processor is configured to identify a first field value of the target field and a first confidence level corresponding to the first field value based on the image data. The processor is further configured to identify a second field value of the target field and a second confidence level corresponding to the second field value based on the first voice data; The processor is further configured to generate a third field value for the target field based on the first field value, the first confidence level, the second field value, and the second confidence level, wherein the third field value is the field value of the target field of the object to be identified.

11. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the object recognition method as described in any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the object recognition method as described in any one of claims 1 to 8.

13. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the object recognition method as described in any one of claims 1 to 8.