Contact information generation method and head-mounted display device
By using image acquisition and voice input devices of a head-mounted display device, combined with face detection and voice processing, contact information is automatically recognized and generated, solving the problems of inconvenience and poor user experience in existing technologies, and realizing convenient contact information recording and generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU LINGBAN TECH CO LTD
- Filing Date
- 2025-04-14
- Publication Date
- 2026-04-24
AI Technical Summary
In existing technologies, when users generate contact information in their address book through the voice assistant of a smart terminal or by manually inputting it, they need to remember or record the contact information, resulting in low convenience and a poor user experience.
By employing a head-mounted display device combined with image acquisition and voice input devices, and through face detection and voice information processing, it automatically identifies and generates contact information, including face feature matching and voice information extraction, to automatically generate contact information.
It improves the convenience of recording contact information, enhances the user experience, and eliminates the need for users to memorize or input information, thus achieving natural information collection and generation.
Smart Images

Figure CN120301963B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to a method for generating contact information and a head-mounted display device. Background Technology
[0002] The address book in a smart device is used to record contact information associated with the user, enabling timely communication with those contacts when needed. Currently, smart devices typically generate contact information in the address book by having the user first obtain the contact's information, and then having the user enter the contact information via the smart device's voice assistant or manually.
[0003] However, in practice, when using the above method to generate contact information in the address book, the following technical problems often occur: users need to memorize or record contact information before they can input it through the voice assistant of the smart terminal or by manually entering it, which makes it less convenient to record contact information and thus results in a poor user experience.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion that follows. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure provide a contact information generation method and a head-mounted display device to address the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide a method for generating contact information, applied to a head-mounted display device, wherein the head-mounted display device includes an image acquisition device and a voice input device. The method includes: in response to detecting confirmed face acquisition information, controlling the image acquisition device to acquire a real-world scene image, and controlling the voice input device to acquire voice information; performing face detection processing on the real-world scene image to obtain at least one face detection information, wherein the face detection information in the at least one face detection information includes face feature information; determining face feature information to be identified based on each face feature information included in the at least one face detection information; performing face matching processing on the face feature information to be identified and a preset face feature information set to obtain matching result information; determining first contact information based on the matching result information; performing information extraction processing on each acquired voice information to obtain second contact information; and generating contact information based on the face feature information to be identified, the first contact information, and the second contact information.
[0008] Optionally, the face detection information in the at least one face detection information further includes a face image, and the determination of the first contact information based on the matching result information includes: in response to determining that the matching result information meets the preset non-matching condition and detecting the new contact confirmation information, determining the face image corresponding to the face feature information to be identified in the at least one face detection information as the contact image; and determining the first contact information based on the contact image.
[0009] Optionally, determining the first contact information based on the contact image includes: performing a filter process on the contact image to obtain a contact filter image; and determining the contact filter image as the first contact information.
[0010] Optionally, the face detection information in the at least one face detection information further includes face bounding box location information, and the determination of the face feature information to be identified based on the face feature information included in the at least one face detection information includes: in response to determining that the number of each face detection information included in the at least one face detection information meets a preset quantity condition, determining the face bounding box location information that meets the preset location condition among the face bounding box location information included in the at least one face detection information as the target face bounding box location information; and determining the face feature information corresponding to the target face bounding box location information among the face feature information included in the at least one face detection information as the face feature information to be identified.
[0011] Optionally, the face detection information in the at least one face detection information further includes a face bounding box number, and the determination of the face feature information to be identified based on the face feature information included in the at least one face detection information includes: in response to detecting image confirmation information, determining the face feature information corresponding to the image confirmation information among the face feature information included in the at least one face detection information as the face feature information to be identified, wherein the image confirmation information includes a confirmed face bounding box number.
[0012] Optionally, the above-mentioned information extraction processing of each collected voice information to obtain the second contact information includes: performing text conversion processing on each collected voice information to obtain text information; determining the above text information as initial text information, and performing the following update steps: displaying the above initial text information in the text preview interface of the display screen of the above-mentioned head-mounted display device, and starting a timing operation; in response to detecting updated voice information corresponding to the initial text information, and not detecting timing end information corresponding to the above-mentioned timing operation, updating the initial text information according to the above-mentioned updated voice information to obtain current text information, and determining the above current text information as initial text information to continue performing the above-mentioned update steps; in response to detecting the above-mentioned timing end information, determining the above current text information as the second contact information.
[0013] Optionally, the aforementioned head-mounted display device further includes a positioning device, and the aforementioned information extraction processing of the collected voice information to obtain second contact information includes: controlling the positioning device to collect user location information; obtaining user current travel information and contact addition serial number; in response to determining that the contact addition serial number meets a preset serial number condition, generating contact-related event information based on the user location information and the user current travel information, and storing the contact-related event information in a preset related event information cache; performing text conversion processing on the collected voice information to obtain voice-text information; and generating second contact information based on the contact-related event information, the voice-text information, and a preset structured contact information slot group.
[0014] Optionally, the above method further includes: in response to determining that the newly added contact number does not meet the preset number condition, and that the user location information meets the preset location condition, and that the current time meets the preset travel condition, obtaining contact-related event information from the preset associated event information cache.
[0015] Optionally, the above method further includes: displaying the contact information on the display screen of the head-mounted display device.
[0016] Optionally, the face detection information in the above-mentioned at least one face detection information further includes face bounding box information, and the above method further includes: displaying the face bounding box information and the contact information in the display screen of the above-mentioned head-mounted display device, wherein the display area of the contact information and the display area of the face bounding box information do not overlap.
[0017] Secondly, some embodiments of this disclosure provide a head-mounted display device, which includes: one or more processors; a storage device storing one or more programs thereon; a display screen for displaying contact information; an image acquisition device for acquiring images of a real-world scene; and a voice input device for acquiring voice information. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0018] The above-described embodiments of this disclosure have the following beneficial effects: the contact information generation method of some embodiments of this disclosure can improve the convenience of recording contact information and enhance the user experience. Specifically, the reason for the low convenience of recording contact information and the poor user experience is that users need to memorize or record contact information before they can input it through the voice assistant of a smart terminal or by manually entering it, resulting in low convenience of recording contact information and thus a poor user experience. Based on this, the contact information generation method of some embodiments of this disclosure is applied to a head-mounted display device, wherein the head-mounted display device includes an image acquisition device and a voice input device. First, in response to detecting and confirming face acquisition information, the image acquisition device is controlled to acquire a real-world scene image, and the voice input device is controlled to acquire voice information. Thus, the real-world scene image and the voice of the user communicating with others in the real-world scene can be obtained in a very natural way without the need for additional user operation, thereby improving the convenience of information acquisition for the user. Then, face detection processing is performed on the real-world scene image to obtain at least one face detection information. The face detection information in the at least one face detection information includes face feature information. This allows us to obtain facial features in real-world scenarios, which can then be used to retrieve contact information from an existing address book. Next, based on the facial feature information included in at least one of the aforementioned facial detection information, the facial features to be identified are determined. This yields the facial features the user needs to query. Then, face matching processing is performed on the facial features to be identified and a preset set of facial features to obtain matching results. This allows us to check if the desired contact information exists in the existing address book, thus identifying the contact. Following this, based on the matching results, the first contact information is determined. This allows us to obtain partial contact information through the query results. Next, information extraction processing is performed on the collected voice information to obtain the second contact information. This allows us to expand, supplement, or update contact information based on voice communication in real-world scenarios, eliminating the need for users to manually add contact information based on the communication content. Finally, based on the facial features to be identified, the first contact information, and the second contact information, contact information is generated. This allows for automatic generation of contact information, eliminating the need for additional user input. Because when generating contact information, the head-mounted display device can collect the user's scene and actual communication content in a very natural way, thereby determining the contacts the user needs and automatically generating contact information in real time. This eliminates the need for the user to remember the content of the communication or perform additional information input operations, thus improving the convenience of recording contact information and enhancing the user experience. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0020] Figure 1 This is a schematic diagram illustrating an application scenario of a contact information generation method according to some embodiments of the present disclosure;
[0021] Figure 2 This is a flowchart of some embodiments of the contact information generation method according to this disclosure;
[0022] Figure 3 These are flowcharts of other embodiments of the contact information generation method according to this disclosure;
[0023] Figure 4 This is a schematic diagram of an application scenario of a contact information generation method according to some embodiments of the present disclosure;
[0024] Figure 5 This is a flowchart of yet another embodiment of the contact information generation method according to the present disclosure;
[0025] Figure 6 This is a schematic diagram of the structure of some embodiments of the head-mounted display device according to the present disclosure. Detailed Implementation
[0026] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0027] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0031] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0032] Figure 1 This is a schematic diagram illustrating an application scenario of a contact information generation method according to some embodiments of the present disclosure.
[0033] exist Figure 1 In the application scenario, firstly, the head-mounted display device 100 can, in response to detecting and confirming face acquisition information 101, control the image acquisition device to acquire a real-world scene image 102, and control the voice input device to acquire voice information 103. Secondly, the head-mounted display device 100 can perform face detection processing on the real-world scene image 102 to obtain at least one face detection information. The face detection information in the at least one face detection information includes face feature information 104. Then, the head-mounted display device 100 can determine the face feature information to be identified 105 based on the face feature information 104 included in the at least one face detection information. Afterwards, the head-mounted display device 100 can perform face matching processing on the face feature information to be identified 105 and a preset face feature information set 106 to obtain matching result information 107. Next, the head-mounted display device 100 can determine the first contact information 108 based on the matching result information 107. Next, the head-mounted display device 100 can extract and process the collected voice information 103 to obtain the second contact information 109. Finally, based on the aforementioned facial feature information 105 to be identified, the aforementioned first contact information 108, and the aforementioned second contact information 109, contact information 110 is generated.
[0034] It should be understood that Figure 1 The number of head-mounted display devices shown is merely illustrative. Any number of head-mounted display devices can be used depending on the implementation requirements.
[0035] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of a contact information generation method according to the present disclosure. This contact information generation method, applied to a head-mounted display device, includes the following steps:
[0036] Step 201: In response to detecting confirmed face acquisition information, control the image acquisition device to acquire real-world scene images and control the voice input device to acquire voice information.
[0037] In some embodiments, the entity executing the contact information generation method (e.g.) Figure 1 The head-mounted display device 100 shown can, in response to detecting and confirming facial recognition information, control an image acquisition device to acquire real-world scene images and control a voice input device to acquire voice information. The head-mounted display device can be a head-mounted device used to assist users in viewing virtual scenes. The head-mounted display device can include, but is not limited to, any of the following: head-mounted augmented display device, head-mounted hybrid display device, or head-mounted virtual reality display device. For example, the head-mounted augmented display device can be AR glasses, and the head-mounted virtual reality display device can be VR glasses. The head-mounted hybrid display device can be MR glasses. The head-mounted display device can include, but is not limited to, an image acquisition device and a voice input device. The image acquisition device can be a device for acquiring images. For example, the image acquisition device can be a camera. The voice input device can be a device for acquiring voice. For example, the voice input device can be a microphone array based on directional noise reduction. The real-world scene image can be an image representing any real-world scene. The voice information can represent a voice signal. The confirmed facial recognition information can be information indicating that the contact person agrees to information collection and that the user confirms the acquisition of a facial image. For example, the aforementioned confirmation of face capture information could be "agree". This confirmation can be triggered by a user via voice command or by a user input operation. The input operation can be triggered by the user pressing a physical button or by selecting an interface control. As an example, the confirmation of face capture information could be triggered by the voice command "start camera".
[0038] Step 202: Perform face detection processing on the real-world scene image to obtain at least one face detection information.
[0039] In some embodiments, the execution entity may perform face detection processing on the real-world scene image to obtain at least one face detection information. The face detection information in the at least one face detection information corresponds to a face image. The face detection information in the at least one face detection information can characterize the information of the detected face image. The face detection information in the at least one face detection information can characterize the information of the detected face image. The face detection information in the at least one face detection information may include, but is not limited to, face feature information. The face feature information can characterize the feature vector of the corresponding face image. In practice, the execution entity may perform face detection processing on the real-world scene image using a preset face detection algorithm to obtain at least one face detection information. The preset face detection algorithm can be a pre-defined algorithm for face detection. For example, the preset face detection algorithm can be a deep learning-based face detection algorithm.
[0040] It should be noted that after performing face detection processing on the aforementioned real-world scene images to obtain at least one face detection information, the aforementioned execution entity can also perform tracking processing on each detected face image using a preset face tracking algorithm. This preset face tracking algorithm can be a feature-point based face tracking algorithm or a deep learning-based face tracking algorithm.
[0041] Step 203: Determine the facial feature information to be identified based on the facial feature information included in at least one face detection information.
[0042] In some embodiments, the execution entity may determine the facial feature information to be identified based on the various facial feature information included in the at least one face detection information. The facial feature information to be identified may be the facial feature information corresponding to the face image that the user needs to identify. In practice, the execution entity may determine the facial feature information included in the at least one face detection information as the facial feature information to be identified in response to determining that the number of face detection information included in the at least one face detection information is 1.
[0043] Optionally, the face detection information in the at least one face detection information may further include face bounding box location information. The face bounding box location information may be the position of a marker box used to mark the corresponding face image on the display screen of the head-mounted display device. The shape of the marker box may be rectangular. The face bounding box location information may include, but is not limited to, the coordinates of the face bounding box center point and the face bounding box size. The face bounding box center point coordinates may be the coordinates of the center point of the marker box on the display screen. The face bounding box size may be the size of the marker box. The face bounding box size can be represented in the form of length * width.
[0044] In some optional implementations of certain embodiments, the execution entity may determine the facial feature information to be identified based on the facial feature information included in the at least one facial detection information through the following steps:
[0045] The first step involves determining the target face frame location information based on the number of face detection pieces included in the at least one face detection information, provided that the number of each face detection piece satisfies a preset quantity condition. Specifically, the preset quantity condition can be that the number of each face detection piece included in the at least one face detection information is greater than 1. The preset position condition can be that the distance between the center point coordinates of the face frame included in the face frame location information and the center point coordinates of the display screen is at its minimum, and the size of the face frame included in the face frame location information is greater than a preset face frame size threshold. This preset face frame size threshold can be a pre-set minimum face frame size. Therefore, smaller detected face images can be eliminated.
[0046] The second step is to determine the facial feature information corresponding to the target face bounding box position information from the facial feature information included in at least one of the above-mentioned facial detection information as the facial feature information to be identified.
[0047] Optionally, the face detection information in the at least one face detection information may further include face bounding box numbers. These face bounding box numbers may be the sequence numbers of the bounding boxes used to mark the corresponding face images. The face bounding box numbers may be sequentially ordered according to the ascending order of the distance of the bounding boxes from the center point of the display screen.
[0048] In some alternative implementations of certain embodiments, the execution entity may, in response to detecting image confirmation information, determine the facial feature information corresponding to the image confirmation information among the various facial feature information included in the at least one face detection information as the facial feature information to be identified. The image confirmation information may represent confirmation of selecting the face image to be identified. The image confirmation information may include, but is not limited to, confirming a face bounding box number. The confirming face bounding box number may be a face bounding box number confirmed by the user. The image confirmation information may be triggered by a user's voice command or generated by the user through input operations.
[0049] As an example, the above image confirmation information can be generated through the following steps:
[0050] The first step involves displaying the marker boxes and face box numbers for each face detection information item on the display screen of the aforementioned head-mounted display device. The face box number is displayed in the upper right corner of the corresponding marker box. The display screen may also include a face box confirmation control. This face box confirmation control can be a control used to confirm the selection of a marker box.
[0051] The second step involves detecting a selection operation applied to any of the aforementioned marker boxes, and determining the face frame number corresponding to that arbitrary marker box as the confirmed face frame number. The selection operation may include, but is not limited to, clicking, swiping, hovering, and gestures.
[0052] Third, in response to the detected selection operation applied to the face frame confirmation control, the face frame number is entered into a preset image confirmation information template to obtain image confirmation information. The preset image confirmation information template can be a pre-defined template used to generate image confirmation information. For example, the preset image confirmation information template could be "Recognize frame A". A can represent the face frame number.
[0053] Therefore, users can choose which contacts to add based on their facial images.
[0054] Step 204: Perform face matching processing on the face feature information to be identified and the preset face feature information set to obtain matching result information.
[0055] In some embodiments, the execution entity can perform face matching processing on the face feature information to be identified and a preset face feature information set to obtain matching result information. The matching result information indicates whether the face match is successful. The preset face feature information in the preset face feature information set can be face feature information pre-stored by the user in a face feature information database. The face feature information database can be a database used to store face feature information. In practice, firstly, for each preset face feature information in the preset face feature information set, the execution entity can determine the similarity between the preset face feature information and the face feature information to be identified as a target similarity. The similarity can be cosine similarity. Then, in response to determining that each of the determined target similarities is less than a preset similarity, preset matching failure information is determined as matching result information. The preset similarity can be a pre-set similarity representing the same image. The preset matching failure information can be pre-set information representing that no matching face image was found. For example, the preset matching failure information can be "0". Subsequently, in response to determining that at least one of the determined target similarities is greater than a preset similarity, the maximum value among the target similarities greater than the preset similarity is determined as the matching similarity. Finally, the matching similarity and the preset matching success information are determined as the matching result information. The preset matching success information can be a pre-defined representation of matching the same face image. For example, the preset matching success information can be "1".
[0056] Step 205: Determine the first contact information based on the matching results.
[0057] In some embodiments, the execution entity can determine the first contact information based on the matching result information. In practice, firstly, in response to determining that the matching result information indicates a successful match, the execution entity can determine the face image corresponding to the matching similarity included in the matching result information as the contact image. Then, it retrieves the preset contact information corresponding to the contact image from the database via a wired or wireless connection, and determines the preset contact information as the first contact information. The preset contact information can be contact information pre-stored by the user. The preset contact information can include, but is not limited to, preset face images. The preset face image can be a pre-set face image. The preset contact information corresponding to the contact image can be: preset contact information including preset face images that are the same as the contact image. The preset contact information can also include a contact identifier and contact information. The contact identifier can be a unique identifier for the contact. For example, the contact identifier can be the contact's name. The contact information can be the contact's contact details. For example, the contact information can be the contact's mobile phone number.
[0058] Optionally, the face detection information in the above-mentioned at least one face detection information may further include a face image. The face image can represent the detected face.
[0059] In some optional implementations of certain embodiments, the aforementioned executing entity may determine the first contact information based on the aforementioned matching result information through the following steps:
[0060] The first step involves determining that the matching result information meets a preset non-matching condition and detecting a new contact confirmation message. The face image corresponding to the facial feature information to be identified in at least one face detection message is then identified as the contact image. The preset non-matching condition can be that the matching result information indicates no matching of the same face image. The new contact confirmation message can indicate that the user confirms the need to add a new contact. This confirmation message can be triggered by the user via voice command, or by an interface control or physical button. For example, the new contact confirmation message could be "Create a new contact".
[0061] The second step is to determine the first contact information based on the aforementioned contact image. In practice, the executing entity can determine the aforementioned contact image as the first contact information.
[0062] In some optional implementations of certain embodiments, the execution entity may determine the first contact information based on the contact image through the following steps:
[0063] The first step is to apply a filter to the contact image to obtain a filtered contact image. In practice, the executing entity can call a preset filter algorithm corresponding to the preset filter type to apply the filter to the contact image and obtain a filtered contact image. The preset filter type can be the system default filter type. The preset filter algorithm can be a pre-defined algorithm for filtering images. Each filter algorithm corresponds to a filter type. The filter types can be, but are not limited to: cartoon style, elderly style, and business style.
[0064] Therefore, users can choose the style of the image according to their own needs, thereby improving the user experience.
[0065] Optionally, before applying a filter to the contact image to obtain a filtered contact image, the executing entity may crop the contact image to obtain a cropped contact image as the contact image.
[0066] Optionally, the aforementioned executing entity may further perform filter processing on the contact image through the following steps to obtain a filtered contact image:
[0067] Step one: In response to detecting user filter selection voice information, the filter identifier included in the user filter selection voice information is determined as the target filter identifier. The filter identifier can be a unique identifier for the corresponding filter type. The user filter selection voice information can represent the user's voice command to select a filter. For example, the user filter selection voice information could be the voice message "Set the face image to cartoon style".
[0068] Step 2: Call the preset filter algorithm corresponding to the target filter identifier to perform filter processing on the contact image to obtain the contact filter image.
[0069] Optionally, the aforementioned executing entity may further perform filter processing on the contact image through the following steps to obtain a filtered contact image:
[0070] Step one: Display the filter confirmation control and corresponding filter identifier controls for each preset filter type on the screen of the aforementioned head-mounted display device. The preset filter types and filter identifier controls can correspond one-to-one. The preset filter types can be pre-defined filter types. The filter identifier controls are for selecting filter identifiers. The filter confirmation control is for confirming the filter style.
[0071] Step 2: In response to the detection of a selection operation on any filter identifier control included in the above-mentioned filter identifier controls, and the detection of a selection operation on the above-mentioned filter confirmation control, the filter identifier corresponding to the above-mentioned arbitrary filter identifier control is determined as the target filter identifier.
[0072] Step 2: Call the preset filter algorithm corresponding to the target filter identifier to perform filter processing on the contact image to obtain the contact filter image.
[0073] The second step is to identify the above-mentioned contact filter image as the first contact information.
[0074] Step 206: Extract information from each of the collected voice messages to obtain the second contact information.
[0075] In some embodiments, the execution entity can perform information extraction processing on the collected voice information to obtain second contact information. This second contact information can be contact information extracted from interactions between the user and a contact in a real-world scenario. In practice, the execution entity can call a preset speech recognition interface to perform information extraction processing on the collected voice information to obtain the second contact information. This preset speech recognition interface can be a pre-configured interface for speech recognition based on a large language model.
[0076] In some optional implementations of certain embodiments, the aforementioned executing entity can perform information extraction processing on the collected voice information to obtain the second contact information through the following steps:
[0077] The first step is to perform text conversion processing on the collected voice information to obtain text information. In practice, the aforementioned execution entity can call a preset speech recognition interface to extract information from the collected voice information and obtain text information. It should be noted that the above text information can be text composed of character combinations, or it can be composed of the text of each slot corresponding to a preset slot label. Each preset slot label corresponds to one slot text.
[0078] The second step is to determine the above text information as the initial text information and perform the following update steps:
[0079] The first update step involves displaying the initial text information and initiating a timing operation on the text preview interface of the head-mounted display device's screen. This text preview interface can be a user-viewable interface for viewing the text. It may also display a timer module and a text confirmation interface. The timer module can be a page module used for timing. The text confirmation interface can be an interface for displaying the text confirmed by the user.
[0080] The second update step, in response to the detection of updated voice information corresponding to the initial text information and the absence of timeout information corresponding to the aforementioned timeout operation, updates the initial text information based on the updated voice information to obtain the current text information, and determines the current text information as the initial text information to continue executing the update step. The updated voice information can be a voice representing the text that needs updating. For example, if the initial text information is "Name Chen Haixin", the updated voice information could be a voice representing "No, Xin should be changed to the three golds, Chen Haixin". The timeout information can represent the interval between the current time and the start of the timeout operation as a preset timeout duration. The preset timeout duration can be a pre-set duration for which the user can correct the text. For example, the preset timeout duration can be 30 seconds. For example, the timeout information can be "end". In practice, the executing entity can continue to call a preset voice recognition interface to update the initial text information based on the updated voice information to obtain the current text information.
[0081] The third update step involves identifying the current text information as the second contact information in response to the detection of the aforementioned timeout information.
[0082] Optionally, the aforementioned executing entity may display the aforementioned second contact information in the aforementioned text confirmation interface.
[0083] Therefore, users can use their voice to correct typos in the text obtained from speech recognition, thereby saving the correct contact information and improving the user experience.
[0084] Furthermore, in the process of adopting the technical solution to solve the technical problems mentioned in the background technology, the following technical problem two further exists: When users are in noisy environments (e.g., exhibitions, factory workshops), the accuracy of speech recognition is low, resulting in low accuracy of the text information obtained by the user, which in turn causes the user to correct the text more frequently, leading to a poor user experience. Therefore, considering the user experience in actual use, the following solution is adopted:
[0085] In some optional implementations of certain embodiments, the aforementioned executing entity can perform information extraction processing on the collected voice information to obtain the second contact information through the following steps:
[0086] The first step is to perform the following sub-steps for each piece of voice information collected:
[0087] Sub-step one involves performing a first speech enhancement process on the aforementioned speech information to obtain first enhanced speech information. In practice, this process can be performed using a preset speech enhancement algorithm to obtain the first enhanced speech information. The preset speech enhancement algorithm can be a pre-defined algorithm for enhancing speech. For example, the preset speech enhancement algorithm could be the GenSE (Generative Speech Enhancement via Language Models using Hierarchical Modeling) algorithm.
[0088] Sub-step two involves performing non-stationary noise enhancement processing on the aforementioned speech information to obtain the second enhanced speech information. In practice, the executing entity can use a preset non-stationary noise enhancement algorithm to perform non-stationary noise enhancement processing on the aforementioned speech information to obtain the second enhanced speech information. The preset non-stationary noise enhancement algorithm can be a pre-defined algorithm used for enhancing speech with non-stationary noise. For example, the preset non-stationary noise enhancement algorithm can be a Kalman filter-based speech enhancement algorithm.
[0089] Sub-step three involves performing scene recognition processing on the aforementioned real-world scene images to obtain the scene type. The scene type characterizes the type of noise environment within the scene. The scene type can be classified as high stationary noise, high non-stationary noise, or low overall noise. For example, a scene with high stationary noise could be a factory workshop. A scene with high non-stationary noise could be an outdoor exhibition. A scene with low overall noise could be a conference room. In practice, the executing entity can use a preset scene classification algorithm to perform scene recognition processing on the aforementioned real-world scene images to obtain the scene type. This preset scene classification algorithm can be a classification algorithm that categorizes scene types. For example, the preset scene classification algorithm could be a convolutional neural network trained on a sample set, taking real-world scene images as input and scene type as output.
[0090] Sub-step four: Based on the aforementioned scenario type, determine the first weight value corresponding to the first enhanced speech information and the second weight value corresponding to the second enhanced speech information. In practice, the executing entity can set the preset first weight value corresponding to the scenario type as the first weight value corresponding to the first enhanced speech information, and set the preset second weight value corresponding to the scenario type as the second weight value corresponding to the second enhanced speech information. The preset first weight value can be a pre-set weight value representing the first enhanced speech information in the final enhanced speech information. The preset second weight value can be a pre-set weight value representing the second enhanced speech information in the final enhanced speech information. It should be noted that the sum of the preset first weight value and the preset second weight value is 1. The higher the non-stationary noise represented by the scenario type, the larger the preset second weight value.
[0091] Sub-step five involves fusing the first enhanced speech information and the second enhanced speech information according to the first weight value and the second weight value to obtain enhanced speech information. In practice, the executing entity can perform weighted fusing processing on the first enhanced speech information and the second enhanced speech information according to the first weight value and the second weight value to obtain enhanced speech information.
[0092] Sub-step six involves performing text recognition processing on the enhanced speech information to obtain speech-text information. This process includes:
[0093] Step one involves performing feature extraction processing on the enhanced speech information to obtain speech feature information. In practice, the executing entity can use a preset speech feature extraction algorithm to perform feature extraction processing on the enhanced speech information to obtain speech feature information. The preset speech feature extraction algorithm can be a pre-defined algorithm for extracting speech features. For example, the preset speech feature extraction algorithm can be the Mel-frequency cepstral coefficient (MFCC) algorithm.
[0094] Step two involves inputting the aforementioned speech feature information into a pre-trained acoustic text probability information generation model to obtain an acoustic text probability information set. This acoustic text probability information set can include both acoustic text and acoustic text probabilities. The acoustic text probability information generation model can be an acoustic model that takes speech feature information as input and outputs the acoustic text probability information set. This acoustic model can be an HMM (Hidden Markov Model)-DNN (Deep Neural Network) model. The acoustic text can be the text generated by the acoustic model. The acoustic text probabilities can be the probabilities of the corresponding acoustic text generated by the acoustic model.
[0095] Step three involves inputting the aforementioned speech feature information into a pre-trained language text generation model to obtain a language text probability information set. This set includes both the language text itself and its probability. The language text generation model can be a language model that takes speech feature information as input and outputs the language text probability information set. This language model can be a Transformer-based language model. The language text can be the text generated by the language model. The language text probability can be the probability of the corresponding language text generated by the language model.
[0096] Step four, for each acoustic text probability information included in the above acoustic text probability information group, perform the following sub-steps:
[0097] The first sub-step involves performing similarity processing on the acoustic text included in the aforementioned acoustic text probability information and the preset dictionary text set to obtain various first similarity scores. The preset dictionary texts in the aforementioned preset dictionary text set can be texts from a pre-defined dictionary containing names, companies, and job titles. It should be noted that user-added names, companies, and job titles can be added to the preset dictionary text set in real time. In practice, for each preset dictionary text included in the preset dictionary text set, the executing entity can determine the first similarity score as the similarity between the acoustic text included in the aforementioned acoustic text probability information and the aforementioned preset dictionary text. This similarity score can be a cosine similarity.
[0098] The second sub-step involves determining the target first similarity score as the maximum value among the first similarities, in response to the determination that each of the aforementioned first similarities satisfies a preset first similarity condition. The preset first similarity condition can be that there exists a first similarity score among the first similarities that is greater than a preset text similarity score. The preset text similarity score can be a similarity score representing that the texts have the same pronunciation.
[0099] The third sub-step is to determine the preset dictionary texts in the preset dictionary text set that correspond to the first similarity of the target as the first target dictionary text.
[0100] The fourth sub-step involves updating the acoustic text included in the above acoustic text probability information to the above first target dictionary text, thereby obtaining updated acoustic text probability information.
[0101] Step 5: For each piece of language text probability information included in the above language text probability information group, perform the following sub-steps:
[0102] The first sub-step involves performing similarity processing on the language texts included in the aforementioned language text probability information and the preset dictionary text set to obtain various second similarities. In practice, for each preset dictionary text included in the preset dictionary text set, the executing entity can determine the similarity between the language texts included in the aforementioned language text probability information and the preset dictionary texts as the second similarity. This similarity can be a cosine similarity.
[0103] The second sub-step involves determining the target second similarity based on the fact that each of the aforementioned second similarities satisfies a preset second similarity condition. The preset second similarity condition can be that there exists a second similarity among the various second similarities that is greater than a preset language-text similarity. The preset language-text similarity can be a pre-defined similarity based on the shared linguistic features of the represented texts.
[0104] The third sub-step involves determining the preset dictionary texts in the preset dictionary text set that correspond to the second similarity to the target as the second target dictionary texts.
[0105] The fourth sub-step involves updating the acoustic text included in the above-mentioned language text probability information to the above-mentioned second target dictionary text, thereby obtaining updated language text probability information.
[0106] Step six: Generate language text information based on the obtained updated acoustic text probability information and updated language text probability information. In practice, the aforementioned execution entity can use a random search algorithm to weight the obtained updated acoustic text probability information and updated language text probability information to obtain the language text information.
[0107] The second step involves inputting the obtained speech-text information into a pre-trained slot label generation model to obtain slot label information. This pre-trained slot label generation model can be a neural network that takes each speech-text information as input and outputs slot label information. This neural network can be a recurrent neural network. The slot label information represents the text in each slot corresponding to each pre-trained slot label. The slot label information can include, but is not limited to, pre-trained slot labels and slot text. There is a one-to-one correspondence between the pre-trained slot labels and the slot text. The pre-trained slot labels can be pre-defined tags for slots. These pre-trained slot labels can include, but are not limited to, job title, name, company, and contact information. The slot text can be the text filled into the corresponding slot.
[0108] The third step is to fill the aforementioned slot label information into the preset contact information slot to obtain the second contact information. The preset contact information slot can be a pre-defined slot for information associated with a contact.
[0109] The above-described technical solution and its related content, as an inventive point of this application, solve technical problem two: "When users are in noisy environments (e.g., exhibitions, factory workshops), the accuracy of speech recognition is low, resulting in low accuracy of the text information obtained by the user, which in turn leads to more corrections by the user and a poor user experience." Factors leading to more text corrections and a poor user experience are often as follows: When users are in noisy environments (e.g., exhibitions, factory workshops), the accuracy of speech recognition is low, resulting in lower accuracy of the text information obtained by the user, which in turn leads to more corrections by the user and a poor user experience. Solving these factors can reduce the number of text corrections by the user and improve the user experience. To achieve this effect, the contact information generation method of this application first performs the following steps on each collected voice information: performing a first speech enhancement process on the voice information to obtain first enhanced voice information; performing non-stationary noise enhancement processing on the voice information to obtain second enhanced voice information; and performing scene recognition processing on the real-world scene image to obtain the scene type. Based on the aforementioned scenario types, a first weight value corresponding to the first enhanced speech information and a second weight value corresponding to the second enhanced speech information are determined. Then, based on the first and second weight values, the first and second enhanced speech information are fused to obtain enhanced speech information. Thus, dynamic speech enhancement processing can be performed on the collected speech information according to the user's scenario, thereby improving the accuracy of speech recognition.Secondly, the enhanced speech information is subjected to text recognition processing to obtain speech-text information. This text recognition processing includes: performing feature extraction on the enhanced speech information to obtain speech feature information; inputting the speech feature information into a pre-trained acoustic text probability information generation model to obtain an acoustic text probability information group, wherein the acoustic text probability information in the acoustic text probability information group includes acoustic text and acoustic text probability; inputting the speech feature information into a pre-trained language text information generation model to obtain a language text probability information group, wherein the language text probability information in the language text probability information group includes language text and language text probability; for each acoustic text probability information group, the following steps are performed: performing similarity processing on the acoustic text included in the acoustic text probability information group and a preset dictionary text set to obtain each first similarity; in response to determining that each first similarity satisfies a preset first similarity condition, the first similarity is... The maximum value among the similarities is determined as the target first similarity; the preset dictionary texts corresponding to the target first similarity in the preset dictionary text set are determined as the first target dictionary texts; the acoustic texts included in the acoustic text probability information are updated to the first target dictionary texts to obtain updated acoustic text probability information; for each language text probability information included in the language text probability information group, the following steps are performed: similarity processing is performed on the language texts included in the language text probability information and the preset dictionary text set to obtain each second similarity; in response to determining that each of the second similarities satisfies the preset second similarity condition, the maximum value among the second similarities is determined as the target second similarity; the preset dictionary texts corresponding to the target second similarity in the preset dictionary text set are determined as the second target dictionary texts; the acoustic texts included in the language text probability information are updated to the second target dictionary texts to obtain updated language text probability information; language text information is generated based on the obtained updated acoustic text probability information and updated language text probability information. Therefore, when recognizing voice information, the generated text can be combined with the information used to generate contact information. By pre-constructing a dictionary, the relevance of the voice-text to the application scenario can be improved, thereby increasing the accuracy of the generated text. Finally, the obtained voice-text information is input into a pre-trained slot label information generation model to obtain slot label information; this slot label information is then filled into preset contact information slots to obtain the second contact information. Thus, contact information can be extracted from voice information through structured information slots.Because when extracting contact information from voice information, combining the noise environment of the user's scene with pre-built dictionary text improves the accuracy of voice-to-text recognition, the accuracy of the generated text can be improved. This reduces the number of times the user needs to correct the text and improves the user experience.
[0110] Step 207: Generate contact information based on the facial feature information of the person to be identified, the first contact information, and the second contact information.
[0111] In some embodiments, the executing entity can generate contact information based on the facial feature information to be identified, the first contact information, and the second contact information. In practice, the executing entity can combine the facial feature information to be identified, the first contact information, and the second contact information to obtain contact information.
[0112] Optionally, the aforementioned implementing entity may store the contact information in a hierarchical and structured manner in various databases.
[0113] As an example, the aforementioned executing entity may store the first contact information, including the contact information, in a direct rendering layer database. This direct rendering layer database can be a database used to store information directly presented during rendering. The second contact information, including the contact information, may be stored in an extended rendering layer database. This extended rendering layer database can be a database used to store information for extended rendering. The facial feature information, including the contact information, may be stored in a facial feature information database.
[0114] Optionally, the aforementioned executing entity may display the contact information on the display screen of the aforementioned head-mounted display device according to a preset presentation level. The aforementioned preset presentation level may be a pre-defined display level on the display screen.
[0115] As an example, the aforementioned executing entity can directly display the first contact information included in the contact information on the display screen of the aforementioned head-mounted display screen, and display the second contact information included in the contact information at the extended display level.
[0116] The above-described embodiments of this disclosure have the following beneficial effects: the contact information generation method of some embodiments of this disclosure can improve the convenience of recording contact information and enhance the user experience. Specifically, the reason for the low convenience of recording contact information and the poor user experience is that users need to memorize or record contact information before they can input it through the voice assistant of a smart terminal or by manually entering it, resulting in low convenience of recording contact information and thus a poor user experience. Based on this, the contact information generation method of some embodiments of this disclosure is applied to a head-mounted display device, wherein the head-mounted display device includes an image acquisition device and a voice input device. First, in response to detecting and confirming face acquisition information, the image acquisition device is controlled to acquire a real-world scene image, and the voice input device is controlled to acquire voice information. Thus, the real-world scene image and the voice of the user communicating with others in the real-world scene can be obtained in a very natural way without the need for additional user operation, thereby improving the convenience of information acquisition for the user. Then, face detection processing is performed on the real-world scene image to obtain at least one face detection information. The face detection information in the at least one face detection information includes face feature information. This allows us to obtain facial features in real-world scenarios, which can then be used to retrieve contact information from an existing address book. Next, based on the facial feature information included in at least one of the aforementioned facial detection information, the facial features to be identified are determined. This yields the facial features the user needs to query. Then, face matching processing is performed on the facial features to be identified and a preset set of facial features to obtain matching results. This allows us to check if the desired contact information exists in the existing address book, thus identifying the contact. Following this, based on the matching results, the first contact information is determined. This allows us to obtain partial contact information through the query results. Next, information extraction processing is performed on the collected voice information to obtain the second contact information. This allows us to expand, supplement, or update contact information based on voice communication in real-world scenarios, eliminating the need for users to manually add contact information based on the communication content. Finally, based on the facial features to be identified, the first contact information, and the second contact information, contact information is generated. This allows for automatic generation of contact information, eliminating the need for additional user input. Because when generating contact information, the head-mounted display device can collect the user's scene and actual communication content in a very natural way, thereby determining the contacts the user needs and automatically generating contact information in real time. This eliminates the need for the user to remember the content of the communication or perform additional information input operations, thus improving the convenience of recording contact information and enhancing the user experience.
[0117] Further reference Figure 3 The diagram illustrates a flow 300 of another embodiment of the contact information generation method according to this disclosure. This contact information generation method, applied to a head-mounted display device, includes the following steps:
[0118] Step 301: In response to detecting confirmed face acquisition information, control the image acquisition device to acquire real-world scene images and control the voice input device to acquire voice information.
[0119] Step 302: Perform face detection processing on the real-world scene image to obtain at least one face detection information.
[0120] Step 303: Determine the facial feature information to be identified based on the facial feature information included in at least one face detection information.
[0121] Step 304: Perform face matching processing on the face feature information to be identified and the preset face feature information set to obtain matching result information.
[0122] Step 305: Determine the first contact information based on the matching results.
[0123] In some embodiments, the specific implementation of steps 301-305 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 201-205 in the corresponding embodiments will not be repeated here.
[0124] Step 306: Control the positioning device to collect user location information.
[0125] In some embodiments, the aforementioned execution entity (e.g. Figure 1 The head-mounted display device 100 shown can control the aforementioned positioning device to collect user location information. The head-mounted display device may further include a positioning device. This positioning device can be used to locate the user's position. For example, the positioning device can be a GPS positioning device. The user location information can represent the user's location. The user location information can include user location coordinates and a user location identifier. The user location coordinates can be the coordinates of the user's location in the GPS positioning coordinate system. The user location identifier can be an identifier of the user's location. For example, the user location identifier can be a building name.
[0126] Optionally, the aforementioned user location information may further include a location scene identifier. This location scene identifier can be a unique identifier of the scene in which the user is located. For example, the location scene identifier can be, but is not limited to, one of the following: classroom, conference room, hotel lobby. The location scene identifier can be obtained by classifying the aforementioned real-world scene image using a preset scene classification algorithm. This preset scene classification algorithm can be a pre-defined algorithm for scene classification. For example, the preset scene classification algorithm can be a support vector machine-based classification algorithm or a neural network-based classification algorithm.
[0127] Step 307: Obtain the user's current itinerary information and the newly added contact number.
[0128] In some embodiments, the aforementioned executing entity can obtain the user's current itinerary information and the newly added contact number. The user's current itinerary information can represent the user's pre-set itinerary for the current time. This information may include, but is not limited to, itinerary time information, itinerary location, and itinerary theme. For example, the itinerary theme could be "xx goods exhibition," where "xx" can be any character. The newly added contact number can be the sequence number of the user's newly added contact within the corresponding itinerary time period. In practice, the aforementioned executing entity can obtain the user's current itinerary information and the newly added contact number from a database via a wired or wireless connection.
[0129] Step 308: In response to determining that the newly added contact number meets the preset number condition, generate contact-related event information based on the user's location information and the user's current trip information, and store the contact-related event information in the preset related event information cache.
[0130] In some embodiments, the executing entity may, in response to determining that the newly added contact number meets a preset sequence number condition, generate contact-related event information based on the user location information and the user's current trip information, and store the contact-related event information in a preset associated event information cache. The preset sequence number condition may be: the newly added contact number is 1. The contact-related event information may represent the contact's activity events. The contact-related event information may include, but is not limited to, trip titles. The preset associated event information cache may be a pre-defined cache for storing contact-related event information. In practice, the executing entity may generate contact-related event information based on the user location information and the user's current trip information using a preset text information extraction algorithm. The preset text information extraction algorithm may be an algorithm for extracting valuable information from text data. For example, the preset text information extraction algorithm may be a template-matching-based information extraction algorithm or an information extraction algorithm based on a Hidden Markov Model (HMM).
[0131] Optionally, the executing entity may also, in response to determining that the newly added contact number does not meet the preset number condition, the user location information meets the preset location condition, and the current time meets the preset travel condition, retrieve contact-related event information from the preset associated event information cache. The preset location condition may be that the user location information is the same as the user location information when the contact was last added. The preset travel condition may be that the current time is within the time period corresponding to the travel time information.
[0132] Therefore, the generated event information can be added to the information of multiple contacts in this trip.
[0133] Step 309: Perform text conversion processing on the collected voice information to obtain voice-text information.
[0134] In some embodiments, the execution entity can perform text conversion processing on the collected speech information to obtain speech-text information. In practice, the execution entity can use a preset speech recognition algorithm to perform text conversion processing on the collected speech information to obtain speech-text information. The preset speech recognition algorithm can be a pre-defined algorithm for converting speech into text. For example, the preset speech recognition algorithm can be a deep learning-based speech recognition algorithm.
[0135] Optionally, the aforementioned execution entity can also perform text conversion processing on the collected voice information by calling a preset speech recognition interface to obtain voice-text information.
[0136] Step 310: Generate second contact information based on contact-related event information, voice and text information, and preset structured contact information slot groups.
[0137] In some embodiments, the executing entity can generate second contact information based on the contact-related event information, the voice / text information, and a preset structured contact information slot group. The preset structured contact information slot group includes preset structured contact information slots that can be slots for tags representing contact information that have been pre-defined. In practice, firstly, the executing entity can combine the contact-related event information and the voice / text information to obtain text information. Then, it performs entity recognition processing on the text information using a preset entity recognition algorithm to obtain an entity tag information group. The preset entity recognition algorithm can be an entity recognition algorithm based on a Conditional Random Field (CRF). The entity tag information in the entity tag information group can include, but is not limited to, entities and entity categories. Finally, for each entity tag information included in the entity tag information group, the following sub-steps are performed:
[0138] The first sub-step involves determining the target slot as the preset structured contact information slot corresponding to the entity tag included in the aforementioned entity tag information within the preset structured contact information slot group. Specifically, the preset structured contact information slot corresponding to the entity tag included in the aforementioned entity tag information can be a preset structured contact information slot whose corresponding slot tag is the same as the entity tag included in the aforementioned entity tag information.
[0139] The second sub-step involves filling the entities included in the aforementioned entity label information into the corresponding target slots.
[0140] It should be noted that when the user's current trip information is empty in the database, contact-related event information can also be extracted from the voice and text information.
[0141] Step 311: Generate contact information based on the facial feature information of the person to be identified, the first contact information, and the second contact information.
[0142] In some embodiments, the specific implementation of step 311 and its resulting technical effects can be found in [reference needed]. Figure 2 Step 207 in the corresponding embodiments will not be repeated here.
[0143] from Figure 3 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 3The process 300 of the contact information generation method in some corresponding embodiments illustrates how to generate extended contact information. Therefore, the solutions described in these embodiments can combine the contact's activity scenarios and events to generate extended contact information, increasing the richness of contact information, thereby improving the user's familiarity with the contact and enhancing the user experience.
[0144] Figure 4 This is a schematic diagram illustrating another application scenario of the contact information generation method according to some embodiments of the present disclosure.
[0145] exist Figure 4 In the application scenario, firstly, the head-mounted display device 100 can, in response to detecting and confirming face acquisition information 101, control the image acquisition device to acquire a real-world scene image 102, and control the voice input device to acquire voice information 103. Secondly, the head-mounted display device 100 can perform face detection processing on the real-world scene image 102 to obtain at least one face detection information. The face detection information in the at least one face detection information includes face feature information 104. Then, the head-mounted display device 100 can determine the face feature information to be identified 105 based on the face feature information 104 included in the at least one face detection information. Afterwards, the head-mounted display device 100 can perform face matching processing on the face feature information to be identified 105 and a preset face feature information set 106 to obtain matching result information 107. Next, the head-mounted display device 100 can determine the first contact information 108 based on the matching result information 107. Next, the head-mounted display device 100 can extract and process the collected voice information 103 to obtain the second contact information 109. Then, based on the aforementioned facial feature information 105, the first contact information 108, and the second contact information 109, contact information 110 is generated. Finally, the head-mounted display device 100 can display the contact information 110 on the display screen.
[0146] It should be understood that Figure 4 The number of head-mounted display devices shown is merely illustrative. Any number of head-mounted display devices can be used depending on the implementation requirements.
[0147] Further reference Figure 5 The diagram illustrates a flow 500 of yet another embodiment of the contact information generation method according to the present disclosure. This contact information generation method, applied to a head-mounted display device, includes the following steps:
[0148] Step 501: In response to detecting confirmed face acquisition information, control the image acquisition device to acquire real-world scene images and control the voice input device to acquire voice information.
[0149] Step 502: Perform face detection processing on the real-world scene image to obtain at least one face detection information.
[0150] Step 503: Determine the facial feature information to be identified based on the facial feature information included in at least one face detection information.
[0151] Step 504: Perform face matching processing on the face feature information to be identified and the preset face feature information set to obtain matching result information.
[0152] Step 505: Determine the first contact information based on the matching results.
[0153] Step 506: Extract information from each of the collected voice messages to obtain the second contact information.
[0154] Step 507 generates contact information based on the facial feature information of the person to be identified, the first contact information, and the second contact information.
[0155] In some embodiments, the specific implementation of steps 501-507 and the resulting technical effects can be referred to steps 201-207 in the embodiments corresponding to 2, and will not be repeated here.
[0156] Step 508: Display the contact information on the display screen of the head-mounted display device.
[0157] In some embodiments, the aforementioned execution entity (e.g. Figure 1 The head-mounted display device 100 can display the contact information on the display screen of the head-mounted display device.
[0158] Optionally, the face detection information in the at least one face detection information may further include face bounding box information. The face bounding box information can represent a face bounding box. The face bounding box can be a bounding box of the detected face image. The face bounding box information may include, but is not limited to, face bounding boxes.
[0159] In some optional implementations of certain embodiments, the executing entity can display face frame information and contact information on the display screen of the head-mounted display device. The display area for the contact information does not overlap with the display area for the face frame information. Therefore, users can view contact information simultaneously while communicating with their contacts without any overlapping, thus improving the user experience.
[0160] from Figure 5 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 5The flow 500 of the contact information generation method in some corresponding embodiments illustrates how to display contact information. Therefore, the solutions described in these embodiments can display the generated contact information in real time on the display screen of a head-mounted display device, making it convenient for users to view during communication with contacts, thereby improving the user experience.
[0161] The following is for reference. Figure 6 It illustrates a head-mounted display device 600 suitable for implementing some embodiments of the present disclosure (e.g., Figure 1 A schematic diagram of the hardware structure of the head-mounted display device 101 in the image. Figure 6 The head-mounted display device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0162] like Figure 6 As shown, the head-mounted display device 600 may include a processing unit 601 (e.g., a central processing unit, a graphics processing unit, etc.), a memory 602, an input unit 603, and an output unit 604. The processing unit 601, memory 602, input unit 603, and output unit 604 are interconnected via a bus 605. Here, the method according to embodiments of this disclosure can be implemented as a computer program and stored in the memory 602. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the method shown in the flowchart. The processing unit 601 in the head-mounted display device implements the voice and text information display method of this disclosure by calling the aforementioned computer program stored in the memory 602. In some implementations, the output unit 604 may include a display screen for displaying contact information. The aforementioned input unit 603 may include an image acquisition device and a voice input device. The image acquisition device may be used to acquire images of a real-world scene. The voice input device may be used to acquire voice information.
[0163] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0164] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0165] The aforementioned computer-readable medium may be included in the aforementioned head-mounted display device; or it may exist independently and not assembled into the head-mounted display device. The aforementioned computer-readable medium carries one or more programs that, when executed by the head-mounted display device, cause the head-mounted display device to: in response to detecting confirmed face acquisition information, control the image acquisition device to acquire a real-world scene image and control the voice input device to acquire voice information; perform face detection processing on the real-world scene image to obtain at least one face detection information, wherein the face detection information in the at least one face detection information includes face feature information; determine the face feature information to be identified based on each face feature information included in the at least one face detection information; perform face matching processing on the face feature information to be identified and a preset set of face feature information to obtain matching result information; determine first contact information based on the matching result information; perform information extraction processing on each acquired voice information to obtain second contact information; and generate contact information based on the face feature information to be identified, the first contact information, and the second contact information.
[0166] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0168] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0169] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for generating contact information, applied to a head-mounted display device, wherein, The head-mounted display device includes an image acquisition device and a voice input device, and the method includes: In response to detecting confirmed face acquisition information, the image acquisition device is controlled to acquire images of the real scene, and the voice input device is controlled to acquire voice information; The real-world scene image is subjected to face detection processing to obtain at least one face detection information, wherein the face detection information in the at least one face detection information includes face feature information; Based on the facial feature information included in the at least one face detection information, determine the facial feature information to be identified; The face feature information to be identified and the preset face feature information set are subjected to face matching processing to obtain matching result information. The preset face feature information in the preset face feature information set is the face feature information that the user has stored in the face feature information database in advance. Based on the matching results, the first contact information is determined; Information extraction processing is performed on each of the collected voice information to obtain the second contact information, wherein the second contact information is the contact information extracted from the communication between the user and the contact in a real-world scenario; Contact information is generated based on the facial feature information to be identified, the first contact information, and the second contact information.
2. The method according to claim 1, wherein, The face detection information in the at least one face detection information further includes a face image, and the determination of the first contact information based on the matching result information includes: In response to determining that the matching result information meets the preset non-matching condition and detecting the new contact confirmation information, the face image corresponding to the face feature information to be identified in the at least one face detection information is determined as the contact image; Based on the contact image, determine the first contact information.
3. The method according to claim 2, wherein, The step of determining the first contact information based on the contact image includes: The contact image is processed with a filter to obtain a filtered contact image; The contact filter image is identified as the first contact information.
4. The method according to claim 1, wherein, The face detection information in the at least one face detection information also includes face bounding box location information, and the step of determining the face feature information to be identified based on each face feature information included in the at least one face detection information includes: In response to determining that the number of each face detection information included in the at least one face detection information meets a preset quantity condition, the face frame position information that meets the preset position condition among the face frame position information included in the at least one face detection information is determined as the target face frame position information. The facial feature information corresponding to the target face bounding box position information among the various facial feature information included in the at least one face detection information is determined as the facial feature information to be identified.
5. The method according to claim 1, wherein, The face detection information in the at least one face detection information also includes a face bounding box number, and the step of determining the face feature information to be identified based on each face feature information included in the at least one face detection information includes: In response to the detection of image confirmation information, the facial feature information corresponding to the image confirmation information among the facial feature information included in the at least one face detection information is determined as the facial feature information to be identified, wherein the image confirmation information includes a confirmed face bounding box number.
6. The method according to claim 1, wherein, The head-mounted display device further includes a positioning device, and the information extraction and processing of the collected voice information to obtain the second contact information includes: The positioning device is controlled to collect user location information; Retrieve user's current itinerary information and the newly added contact number; In response to determining that the newly added contact number meets the preset number condition, the system generates contact-related event information based on the user's location information and the user's current trip information, and stores the contact-related event information in a preset related event information cache. The collected voice information is processed into text to obtain voice-text information; The second contact information is generated based on the contact-related event information, the voice and text information, and the preset structured contact information slot group.
7. The method according to claim 6, wherein, The method further includes: In response to determining that the newly added contact number does not meet the preset number condition, the user location information meets the preset location condition, and the current time meets the preset travel condition, the contact-related event information is retrieved from the preset associated event information cache.
8. The method according to claim 1, wherein, The method further includes: The contact information is displayed on the screen of the head-mounted display device.
9. The method according to claim 8, wherein, The face detection information in the at least one face detection information further includes face bounding box information, and the method further includes: The face frame information and the contact information are displayed on the display screen of the head-mounted display device, wherein the display area of the contact information does not overlap with the display area of the face frame information.
10. The method according to claim 1, wherein, The step of extracting and processing the collected voice information to obtain the second contact information includes: The collected voice information is processed into text information through text conversion. The text information is determined as the initial text information, and the following update steps are performed: The initial text information and the start timing operation are displayed in the text preview interface of the display screen of the head-mounted display device; In response to the detection of updated voice information corresponding to the initial text information and the absence of timing end information corresponding to the timing operation, the initial text information is updated according to the updated voice information to obtain the current text information, and the current text information is determined as the initial text information to continue executing the update step; In response to detecting the end of the timeout information, the current text information is identified as the second contact information.
11. A head-mounted display device, comprising: One or more processors; A storage device on which one or more programs are stored; The display screen is used to show contact information; Image acquisition device, used to acquire images of real-world scenes; A voice input device for collecting voice information; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Face identification-based communication method and apparatus
CN104112119A
Mobile terminal and control method thereof
US20150373174A1