Interaction method and device based on visual earphone, computer equipment and medium

By integrating cameras and beam transmitters in headphones, combining computer vision and natural language processing technology, the problem of single interaction methods of traditional headphones is solved, intelligent environmental information acquisition and personalized feedback are achieved, and user experience is improved.

CN120583181APending Publication Date: 2025-09-02SHENZHEN RB LINK INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510465109.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The interaction method of traditional headphones is single and lacks intelligence, which cannot meet the needs of complex scenarios.

Method used

By integrating the camera and beam transmitter in the headset, using interactive instructions to judge user needs, controlling the camera to scan and process images, and combining computer vision and natural language processing technology to provide personalized feedback.

Benefits of technology

It realizes intelligent interaction between the headset and the user, and can obtain environmental information in real time without relying on external devices, improving the user's practicality and interactive experience in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583181A_ABST
    Figure CN120583181A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of interaction based on visual earphones, and discloses an interaction method and device based on a visual earphone, computer equipment and a medium, and the method comprises the steps: judging whether a user wearing the visual earphone sends an interaction instruction or not; analyzing content in the interaction instruction to determine a user demand; in response to the user demand, controlling a camera to scan current content to obtain a scanned picture; performing data processing on the scanned picture based on the user demand to obtain a processing result; and feeding back the processing result to the user. The method has the beneficial effects that the user demand is determined through the interaction instruction, and intelligent and personalized services can be provided according to the user demand. Through the camera integrated with the earphone, a user can obtain environment information in real time without depending on external equipment, and obtain instant feedback through the earphone, so that the practicability and interaction experience of the intelligent wearable equipment in a complex environment are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of interactive technology based on visualization earphones, and in particular to an interactive method, device, computer equipment and medium based on visualization earphones. Background Art

[0002] With the continuous advancement of intelligent technology, wearable devices are gradually gaining widespread adoption in daily life and various professional applications, becoming an indispensable part of people's lives. Especially in cutting-edge fields such as augmented reality (AR), virtual reality (VR), and intelligent assistive devices, headphones, as a key interactive tool, are gradually taking on more functions. In addition to audio output and voice calls, headphones can also interact with the external environment through sensors, acquiring environmental information and providing real-time feedback to users, enhancing the overall user experience. In some professional scenarios, such as medical, industrial, and entertainment fields, the intelligent functions of headphones have gradually expanded, becoming a key device for providing an immersive experience and improving work efficiency.

[0003] Traditional headphones are mostly limited to audio output or voice call functions. The interaction method between users and devices is single, lacking sufficient intelligence and unable to meet the needs of complex scenarios. Summary of the Invention

[0004] Based on this, it is necessary to address the existing interaction problems based on visualization headphones and propose an interaction method, device, computer equipment and medium based on visualization headphones.

[0005] An interactive method based on a visual headset, the visual headset comprising a main control center, a headset body, and a camera, the headset body and the camera being respectively connected to the main control center, the method comprising:

[0006] Determining whether a user wearing the visualization headset issues an interaction instruction;

[0007] Analyzing the content of the interactive instructions to determine user needs;

[0008] In response to the user's demand, the camera is controlled to scan the current content to obtain a scanned image;

[0009] Performing data processing on the scanned image based on the user's needs to obtain a processing result;

[0010] Feedback the processing result to the user.

[0011] Furthermore, the visualization headset further includes a light beam transmitter, which is connected to the main control center. The step of performing data processing on the scanned image based on the user's needs to obtain a processing result includes:

[0012] Determining whether the user has issued a first instruction;

[0013] If the first instruction is issued, marking the content in the scanned image based on the light beam emitter;

[0014] Determining whether the user has issued a second instruction;

[0015] If the second instruction is issued, stopping marking the content in the scanned image and extracting the marked content;

[0016] Data analysis is performed on the marked content based on the user demand to obtain a processing result.

[0017] Furthermore, if the second instruction is issued, the step of stopping marking the content in the scanned image and extracting the marked content includes:

[0018] Acquire a first position in the scanned image where the light beam emitter emits light after the first instruction is issued, and a second position in the scanned image where the light beam emitter emits light after the second instruction is issued;

[0019] The content from the first position to the second position is marked, and the marked content is extracted.

[0020] Furthermore, the step of analyzing the content of the interaction instruction to determine the user's needs includes:

[0021] Converting the content of the interactive instruction into text information to obtain an interactive text instruction;

[0022] The interactive text instructions are semantically analyzed through a natural language processing model to obtain user needs.

[0023] Furthermore, the step of performing data processing on the scanned image based on the user's needs to obtain a processing result includes:

[0024] Determining associated feature information based on the user needs;

[0025] Utilizing a preset computer vision technology, and based on the feature information, extracting key information from the scanned image;

[0026] The key information is processed based on the user's needs to obtain a processing result.

[0027] Furthermore, the feature information includes at least one of text features, picture features, physical features, and QR code features.

[0028] Furthermore, the visualization headset further includes a bone conduction generator, and the step of feeding back the processing result to the user includes:

[0029] sending the processing result to the bone conduction generator;

[0030] The bone conduction generator converts the processing result into a mechanical signal and transmits it to the user.

[0031] An interactive device based on a visual headset, the visual headset comprising a main control center, a headset body, and a camera, the headset body and the camera being respectively connected to the main control center, the device comprising:

[0032] a judgment module, configured to judge whether a user wearing the visualization headset has issued an interaction instruction;

[0033] An analysis module, configured to analyze the content of the interaction instruction to determine user needs;

[0034] A response module, configured to control the camera to scan the current content in response to the user's demand to obtain a scanned image;

[0035] A processing module, configured to perform data processing on the scanned image based on the user's needs to obtain a processing result;

[0036] A feedback module is used to feed back the processing result to the user.

[0037] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0038] Determining whether a user wearing the visualization headset issues an interaction instruction;

[0039] Analyzing the content of the interactive instructions to determine user needs;

[0040] In response to the user's demand, the camera is controlled to scan the current content to obtain a scanned image;

[0041] Performing data processing on the scanned image based on the user's needs to obtain a processing result;

[0042] Feedback the processing result to the user.

[0043] A computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the following steps:

[0044] Determining whether a user wearing the visualization headset issues an interaction instruction;

[0045] Analyzing the content of the interactive instructions to determine user needs;

[0046] In response to the user's demand, the camera is controlled to scan the current content to obtain a scanned image;

[0047] Performing data processing on the scanned image based on the user's needs to obtain a processing result;

[0048] Feedback the processing result to the user.

[0049] The beneficial effects of this invention include determining user needs through interactive instructions and providing intelligent, personalized services based on these needs. Through the headset's integrated camera, users can obtain real-time environmental information without relying on external devices and receive instant feedback through the headset, greatly enhancing the practicality and interactive experience of smart wearable devices in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0051] in:

[0052] Figure 1 A diagram illustrating an application environment of an interactive method based on a visual headset in one embodiment;

[0053] Figure 2 A schematic diagram of a structure based on a visual headset in one embodiment;

[0054] Figure 3 is a flowchart of interaction based on a visual headset in one embodiment;

[0055] Figure 4 is a structural block diagram of an interactive device based on a visual headset in one embodiment;

[0056] Figure 5 FIG. 1 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0058] Figure 1 FIG is a diagram of an interactive application environment based on a visual headset in one embodiment. Figure 1 This visual headset-based interaction method is applied to a visual headset-based interaction system. This visual headset-based interaction system includes a terminal 110 and a server 120. Terminal 110 and server 120 are connected via a network. Terminal 110 can be a desktop terminal or a mobile terminal. The mobile terminal can be at least one of a mobile phone, tablet computer, and laptop computer. Server 120 can be implemented as a standalone server or a server cluster consisting of multiple servers. Terminal 110 is used to collect and analyze data, and server 120 is used to provide corresponding services to the terminal.

[0059] like Figure 2 As shown, the visual headset includes a main control center (not shown), a headset body 14 and a camera 12, and the headset body 14 and the camera 12 are respectively connected to the main control center. The camera 12 is arranged on the end of the side arm on one side of the headset body 14, and the headset body 14 is also provided with a first rotating part 11, which can drive the camera 12 to rotate. In addition, the camera 12 can also be rotatably connected to the headset body 14 so that the camera 12 can be partially rotated, which is convenient for adjusting the position of the camera 12. Similarly, a light beam emitter 13 is also provided at the end of the side arm on the other side of the headset body 14, and the headset body 14 is also provided with a second rotating part 15. The second rotating part 15 can drive the light beam emitter 13 to rotate for adjusting the position of the light beam emitter 13. The light beam emitter 13 can also be connected to the main control center.

[0060] like Figure 3 As shown, in one embodiment, a method for interactive use based on a visual headset is provided. This method can be applied to both a terminal and a server. This embodiment uses the application to a terminal as an example. The method for interactive use based on a visual headset specifically includes the following steps:

[0061] S1: Determine whether the user wearing the visualization headset issues an interaction instruction;

[0062] S2: Analyze the content of the interactive instruction to determine user needs;

[0063] S3: controlling the camera 12 to scan the current content in response to the user's demand to obtain a scanned image;

[0064] S4: performing data processing on the scanned image based on the user's needs to obtain a processing result;

[0065] S5: Feedback the processing result to the user.

[0066] As described in the above step S1, it is determined whether the user wearing the visual headset has issued an interactive instruction, wherein the interactive instruction can be a gesture instruction, a voice instruction, or a tactile input. For example, the user can complete the payment by inputting a voice instruction, such as synchronous pronunciation of text, interpretation of text, translation of text, interpretation of real objects, explanation of the artistic conception of pictures, and QR code recognition. It can also be a gesture instruction, and different gestures are pre-set in the system to represent different instructions. It can also be tactile input, that is, the visual headset is equipped with a touch-sensing area, and different instruction inputs are realized by inputting corresponding tactile inputs in the touch-sensing area. In a preferred embodiment, voice instructions are used as interactive instructions. The voice instructions contain more information, which is more conducive to the user interacting with the visual headset.

[0067] As described in step S2 above, the content of the interaction instruction is analyzed to determine the user's needs. Specifically, this can be processed using natural language processing. If the interaction content is voice, the voice signal can be converted into text information, which is then analyzed using natural language processing to extract the user's specific needs from the text information. For example, the instruction "Scan the environment and identify objects" may indicate that the user wants the headset to scan the surrounding environment and identify objects within it. If there are multiple rounds of dialogue, the system can also consider all interactions to determine the user's precise needs.

[0068] As described in step S3 above, in response to the user's request, the camera 12 is controlled to scan the current content to obtain a scanned image. The system activates the image capture function by controlling the camera 12 in the headset to begin real-time filming or video recording. Based on the user's request, the camera 12 can scan the entire surrounding environment, capturing still images or videos to obtain visual information. The camera 12 can adjust the scanning area or direction based on user instructions, such as scanning an object or scene in front. The camera 12 can also adjust the angle and focal length to ensure the desired image is captured.

[0069] As described in step S4 above, the scanned image is processed based on the user's needs to obtain a processing result. The content of the data processing varies depending on the user's needs. For example, if the user needs object recognition, computer vision techniques such as convolutional neural networks (CNNs) can be used. Specifically, before object recognition, the image often requires preprocessing, including denoising, resizing, grayscale conversion, and edge detection. These operations can improve recognition accuracy and efficiency. Traditional computer vision methods perform object recognition by extracting features from images. For example, common feature extraction methods include SIFT (Scale Invariant Feature Transform), SURF (Speeded Robust Features), and HOG (Histogram of Oriented Gradients). The goal of feature extraction methods is to extract key points and descriptors from the image so that they can be matched with other images to identify the object. If text recognition is required, if the image contains text (such as street signs or store signs), the system may perform optical character recognition (OCR) to extract the text information. If the user's needs involve scene analysis (such as the layout and color of the surrounding environment), the system will use image processing techniques to analyze these image features and generate corresponding data.

[0070] As described in step S5 above, the processing result is fed back to the user. The feedback method can be diversified according to the function of the device and user needs, such as voice feedback, in which the system feeds back the result to the user in the form of voice through the headphone speaker. For example, the user may hear "a red car is recognized" or "no objects are recognized around". If the visual headset is equipped with a built-in display (such as augmented reality glasses), the processing result may be displayed on the display in the form of an image or text, and the user can view it directly. For some important feedback or reminders, the system may prompt by vibration. According to needs, the system can also guide the user to take the next step, such as "please move left to scan more areas" or "please say more information for further analysis". In this way, intelligent and personalized services can be provided according to user needs. Through the camera 12 integrated in the headset, the user can obtain environmental information in real time without relying on external devices, and obtain instant feedback through the headset, which greatly improves the practicality and interactive experience of smart wearable devices in complex environments.

[0071] This embodiment provides an interactive method based on a visual headset, suitable for user-device interaction in augmented reality (AR) applications. The specific implementation of the visual headset includes a main control center, a headset body, and a camera. The main control center is the core processing unit, the headset body is a device worn on the user's head, and the camera is mounted on the headset body and is responsible for capturing the user's surroundings. In actual application, when a user puts on the visual headset and turns on the device, the system first determines whether the user has issued an interactive command using the headset's built-in sensors and voice recognition module. The user can issue a command through voice, gesture, or other biometric recognition methods. For example, if the user says "scan surrounding objects" through a voice command or performs a specific gesture using gestures, the system will recognize the command and begin executing subsequent steps. Once the user confirms the interactive command, the system analyzes the command content and determines the user's needs. In this example, the user may need to scan a specific object. After analyzing the command, the system controls the headset's camera to automatically focus on the target object. The camera begins scanning the current scene, capturing image or video data, and transmitting it to the main control center for processing. After obtaining the scanned image, the main control center processes the image data based on the user's needs. For example, suppose a user wants to scan and identify surrounding plants. The system uses image recognition technology to analyze the plants in the scanned image, extract their characteristic information, and classify and identify them. The results may include information such as the plant's name, species, and characteristics. Finally, the control center provides feedback to the user via voice, display, or other feedback methods. For example, in a virtual reality application, the user might hear a voice prompt played through the headset, informing them that the scanned image is a "common willow tree" and providing further plant information. The user can then control the device with further commands for more in-depth operation. In this way, the visual headset-based interaction method not only enables intelligent interaction between the device and the user, but also provides real-time environmental data analysis and feedback, greatly improving the user experience. This demonstrates its strong application potential, particularly in scenarios requiring rapid information acquisition, virtual training, or education.

[0072] In one embodiment, the visualization headset further includes a light beam transmitter 13, and the light beam transmitter 13 is connected to the main control center. The step S4 of performing data processing on the scanned image based on the user's needs to obtain a processing result includes:

[0073] S401: Determine whether the user has issued a first instruction;

[0074] S402: If the first instruction is issued, marking the content in the scanned image based on the light beam emitter 13;

[0075] S403: Determine whether the user has issued a second instruction;

[0076] S404: If the second instruction is issued, stop marking the content in the scanned image and extract the marked content;

[0077] S405: Perform data analysis on the marked content based on the user's needs to obtain a processing result.

[0078] As described in step S401 above, it is determined whether the user has issued a first instruction. A instruction can be issued through voice, gesture, or tactile input, the specific method depending on the design of the visualization headset. For example, the user can issue this instruction by saying "Mark the object in the current image" or through gesture control. In addition, it should be noted that the light beam emitter 13 needs to be turned on before issuing the first instruction. Specifically, it can be turned on by the user inputting other instructions, and the content of the instruction is preferably a voice instruction.

[0079] As described in step S402 above, the system uses the beam emitter 13 to mark specific content in the scanned image. The beam emitter 13 may emit a laser or other type of light beam (such as an LED) and focus the beam on a target area in the image (such as an object, a person, or a specific area). When the beam emitter emits a light beam, the camera 12 in the headset starts to capture the scanned image. The camera 12 records the image of the current environment in real time, including the area illuminated by the light beam. At this time, the image captured by the camera will contain the mark formed by the light beam emitter, and the captured image will be transmitted to the processing unit. The system will analyze this image through an image processing algorithm to identify the marked area projected by the light beam. The specific process includes: using an edge detection algorithm (such as Canny edge detection) to identify bright spots or high-contrast areas in the image, extracting feature information of the area illuminated by the light beam, such as color, shape, size, etc., determining the boundary of the area illuminated by the light beam, and generating a specific marked content area.

[0080] As described in step S403 above, it is determined whether the user has issued a second instruction, i.e., requesting to stop the marking process and proceed with subsequent data processing. This step is also completed through voice commands, gesture control, or other input methods. For example, the user can say "stop marking" by voice or use a specific gesture to indicate the end of marking and the start of processing the marked content.

[0081] As described in step S404 above, if the second instruction is issued, the marking of the content in the scanned image is stopped and the marked content is extracted. Once the user issues the second instruction, the system stops the marking process and begins extracting the marked content from the scanned image. At this time, the system turns off the beam emitter 13, stops emitting the beam or laser, and extracts the marked content. The system extracts the image area or object previously marked by the beam for further analysis.

[0082] As described in step S405 above, data analysis is performed on the marked content based on the user's needs to obtain a processing result. The specific data analysis method needs to be determined according to the marked content. For example, if the user marks objects in the image, the system can identify these objects through computer vision algorithms and extract the characteristics or categories of the objects. If the user marks the text area in the image (such as store signs, document content, etc.), the system can perform optical character recognition (OCR) to extract the text information therein. If a specific area in the environment is marked, the system can analyze the characteristics of the area, such as measuring distance, detecting the relationship between the object and the environment, etc.

[0083] In one embodiment, if the second instruction is issued, the step S404 of stopping marking the content in the scanned image and extracting the marked content includes:

[0084] S4041: Acquire a first position in the scanned image where the light beam emitter 13 emits light after the first instruction is issued, and a second position in the scanned image where the light beam emitter 13 emits light after the second instruction is issued;

[0085] S4042: Mark the content from the first position to the second position, and extract the marked content.

[0086] As described in the above steps S4041-S4042, the first position in the scanned image where the light beam emitter 13 emits light after the first instruction is issued, and the second position in the scanned image where the light beam emitter 13 emits light after the second instruction is issued are obtained; the content from the first position to the second position is marked, and the marked content is extracted. Based on the acquired first position and second position, the system will mark the area between these two positions in the image. The marking may be achieved by the light beam emitter 13 emitting a laser or other type of light beam, which is irradiated on a specific area. Through the continuous emission of the light beam, the system can clearly delineate this area, thereby clearly highlighting or marking the area in the image. It should be noted that the light beam emitter 13 can frame the content to be marked by underlining or in the form of a box.

[0087] In one embodiment, step S2 of analyzing the content of the interaction instruction to determine the user's needs includes:

[0088] S201: Convert the content of the interactive instruction into text information to obtain an interactive text instruction;

[0089] S202: Perform semantic analysis on the interactive text instructions through a natural language processing model to obtain user needs.

[0090] As described in steps S201-S202 above, user interaction commands are typically issued via voice. The system uses advanced speech recognition technology to convert the voice signals into processable text. This process typically involves multiple stages. First, the system uses a speech recognition engine to convert the user's voice signals into text data. Next, it uses a natural language processing (NLP) model to conduct in-depth analysis of this text information. Specifically, the system breaks down the complete user command into individual words. This process helps understand the function of each word within the sentence and accurately identifies the relationships between words, such as the different grammatical roles of nouns, verbs, and adjectives. Next, the system further analyzes the structure of the entire sentence, identifying its grammatical framework, understanding the core meaning of the sentence, and ensuring that each part is correctly interpreted. By incorporating machine learning and deep learning models, the system can gradually learn from user interactions and identify the user's specific needs through a trained intent classification model. This classification process can categorize different types of commands into specific intents, such as "scan an object," "query information," or "activate a mode." In this way, the system can accurately understand the user's request and respond accordingly. Semantic analysis relies not only on understanding vocabulary and grammar but also incorporates contextual information to ensure that the system can correctly identify user needs even in complex environments, avoiding misunderstandings or incorrect responses. This precise semantic analysis capability enables the system to provide more intelligent and personalized services, making the visual headset an efficient and accurate interactive tool, greatly improving the user experience.

[0091] In one embodiment, the step S4 of performing data processing on the scanned image based on the user's needs to obtain a processing result includes:

[0092] S411: Determine associated feature information based on the user demand;

[0093] S412: Utilizing a preset computer vision technology and extracting key information from the scanned image based on the feature information;

[0094] S413: Perform data processing on the key information based on the user's needs to obtain a processing result.

[0095] As described in steps S411-S413 above, the system first reviews and deeply understands the user's needs, clarifying the specific type of information the user desires, such as "identifying objects," "extracting text," or "performing scene analysis." This step is crucial as it provides direction and basis for subsequent image analysis. Based on the user's needs, the system lists feature information that may be relevant to the needs. Feature information generally refers to key attributes that can help describe and identify image content, including object shape, color, and texture, as well as text within the image. Accurate feature extraction is fundamental to the image understanding process and can significantly improve the efficiency and accuracy of subsequent processing and analysis. Next, the system selects appropriate computer vision technology based on the identified feature information to process the scanned image. Depending on the needs, the system may employ different technologies. For example, deep learning methods such as convolutional neural networks (CNNs) can be combined with image recognition technology to identify and locate specific objects in images. Common technologies include YOLO (You Only Look Once) and Faster R-CNN, which can efficiently detect and classify objects in real-time scenarios. For tasks requiring precise analysis of image details, the system may use semantic segmentation or instance segmentation techniques, such as Mask R-CNN. This approach helps the system classify images at the pixel level and extract regions of interest. Furthermore, the system can utilize classic feature extraction algorithms, such as SIFT (Scale-Invariant Feature Transform), SURF (Speeded Robust Features), and HOG (Histogram of Oriented Gradients), to extract key features from the image and further analyze them, thereby enhancing the system's understanding of image content. This process not only improves system accuracy but also enables it to flexibly adapt to various complex environments and meet diverse user needs. The system processes the scanned image to extract key information related to the associated features. This process may involve scanning and analyzing the entire image, as well as detecting each feature individually. The extracted key information may include: identified objects, such as "a cat was identified." Extracted text, such as "the road sign reads 'Leading to the city center.'" Scene descriptions, such as "the scene is a park with trees and benches." Data processing is then performed based on the extracted feature information and user needs, resulting in a user-specific result. As for the analysis of feature information, it needs to be processed in combination with the user's needs. Specifically, for example, if the user needs to translate or parse text, it can be processed through the corresponding processing method. The specific processing methods for different needs are described in detail above and will not be repeated here.

[0096] In one embodiment, the feature information includes at least one of text features, image features, physical features, and QR code features.

[0097] In one embodiment, the visualization headset further includes a bone conduction generator, and the step S5 of feeding back the processing result to the user includes:

[0098] S501: Sending the processing result to the bone conduction generator;

[0099] S502: The bone conduction generator converts the processing result into a mechanical signal and transmits it to the user.

[0100] As described in steps S501-S502 above, the system sends the formatted and processed results to the bone conduction generator via the internal communication module, ensuring timely and accurate delivery of the information to the user. After receiving the processed results from the system, the bone conduction generator uses its built-in sound-generating unit (such as a piezoelectric oscillator with a frequency range of 200Hz-1000Hz) to convert these processed results into mechanical vibration signals. A bone conduction generator typically consists of the following components: The vibration unit is the core component, converting electrical signals into mechanical vibrations. These vibrations are typically achieved using piezoelectric materials or electromagnetic drive technology. Piezoelectric materials convert electrical signals into mechanical vibrations through the piezoelectric effect (the deformation of a material under the influence of voltage). Common piezoelectric materials include lead titanate. Electromagnetic drive: Current is passed through a coil to generate a magnetic field, driving the vibrating element to move, thereby achieving sound conversion. Conductive contact components: This component directly contacts the user's skin. Common conductive components include: ear pads or headbands: These are used to securely attach the vibration unit to the user's skull, typically close to the temporal bone. Circuitry and signal processing units: These include receiving, processing, amplifying, and conditioning sound signals. Signals may be transmitted wirelessly (such as via Bluetooth) from a mobile phone or other audio source. The device's internal circuitry converts these signals into signals suitable for bone conduction. This signal is generated by a vibrator and transmitted as high-frequency mechanical vibrations to the user's skull, specifically the temporal bone. Bone conduction technology works by transmitting sound signals to the inner ear through the temporal bone, bypassing the traditional conduction pathways of the outer and middle ears and directly affecting the inner ear's auditory system. Leveraging the unique advantages of bone conduction, users can maintain awareness of surrounding sounds while hearing feedback, without creating a complete soundproofing effect. This allows users to receive feedback without losing sensitivity to their surroundings, making them particularly suitable for use in noisy environments such as streets and construction sites. Furthermore, this technology offers a superior listening experience for the hearing impaired. Because bone conduction technology bypasses the middle and outer ears, transmitting sound directly through the bones, bone conduction headphones can still effectively deliver sound information to users with middle or outer ear damage, helping them better absorb surrounding sounds and information.

[0101] In short, bone conduction-based feedback not only enhances user experience and convenience, but also expands the technology's application, particularly for those with hearing impairments, effectively improving their communication and interaction methods. This allows users to access information more conveniently and safely, while also maintaining awareness of their surroundings in complex living and working environments.

[0102] In one embodiment, a visual headset can integrate a camera, a beam emitter, and bone conduction headphones. The camera can provide images and visual information for environmental awareness, object recognition, navigation, and visual feedback. The beam emitter (such as a laser pointer or projector) is used to mark or highlight specific targets or areas of user interest, which can help users more clearly understand or locate the content they need to focus on. Bone conduction further assists users through audio feedback. Through bone conduction headphones, users can receive voice prompts, instructions, or warnings without being distracted by external noise, while maintaining auditory awareness of the environment. In work, training, or other environments that require concentration, users can use the beam emitter to mark objects or areas that require special attention. For example, in remote collaboration, the real-time image captured by the camera can be used to mark the target object (such as a device or file) using the beam emitter, while the bone conduction headphones can provide relevant voice instructions or additional information to help users focus on important content and avoid misoperation. Combining the real-time data from the beam emitter and the camera, users can more accurately know the objects they are concerned about or operating. The beam transmitter helps users identify and mark targets, while the bone conduction headset provides voice prompts during annotation, making the interaction more natural. Integrating the camera, beam transmitter, and bone conduction headset into a single system can greatly enhance the user's multi-sensory experience, improve operational efficiency, and provide real-time feedback and guidance. This multi-dimensional technological fusion can improve work accuracy, reduce interference, enhance safety, and adapt to a variety of special demand scenarios (such as industry, healthcare, and emergency response).

[0103] Reference Figure 4 The present invention also provides an interactive device based on a visual headset, wherein the visual headset includes a main control center, a headset body, and a camera, wherein the headset body and the camera are respectively connected to the main control center, and the device includes:

[0104] A judgment module 10 is used to judge whether the user wearing the visualization headset has issued an interaction instruction;

[0105] An analysis module 20 is configured to analyze the content of the interaction instruction to determine user needs;

[0106] The response module 30 is used to control the camera to scan the current content in response to the user's demand to obtain a scanned image;

[0107] The processing module 40 is used to perform data processing on the scanned image based on the user's needs to obtain a processing result;

[0108] The feedback module 50 is configured to feed back the processing result to the user.

[0109] In one embodiment, the visualization headset further includes a light beam transmitter connected to the main control center. The processing module 40 includes:

[0110] A first instruction judging submodule, configured to judge whether the user has issued a first instruction;

[0111] a marking submodule, configured to mark the content in the scanned image based on the light beam emitter if the first instruction is issued;

[0112] A second instruction judging submodule, configured to judge whether the user has issued a second instruction;

[0113] a marked content extraction submodule, configured to stop marking the content in the scanned image and extract the marked content if the second instruction is issued;

[0114] The data analysis submodule is used to perform data analysis on the marked content based on the user's needs to obtain a processing result.

[0115] In one embodiment, the tag content extraction submodule includes:

[0116] an acquiring unit, configured to acquire a first position in the scanned image where the light beam emitter emits light after the first instruction is issued, and a second position in the scanned image where the light beam emitter emits light after the second instruction is issued;

[0117] The marking submodule is used to mark the content from the first position to the second position and extract the marked content.

[0118] In one embodiment, the analysis module 20 includes:

[0119] A conversion submodule, configured to convert the content in the interactive instruction into text information to obtain an interactive text instruction;

[0120] The semantic analysis submodule is used to perform semantic analysis on the interactive text instructions through a natural language processing model to obtain user needs.

[0121] In one embodiment, the processing module 40 includes:

[0122] A feature information determination submodule, configured to determine associated feature information based on the user's needs;

[0123] A key information extraction submodule, configured to extract key information from the scanned image based on the feature information using a preset computer vision technology;

[0124] The data processing submodule is used to perform data processing on the key information based on the user's needs to obtain a processing result.

[0125] In one embodiment, the feature information includes at least one of text features, image features, physical features, and QR code features.

[0126] In one embodiment, the visualization headset further includes a bone conduction generator, and the feedback module 50 includes:

[0127] a processing result sending submodule, configured to send the processing result to the bone conduction generator;

[0128] The processing result conversion submodule is used to convert the processing result into a mechanical signal through the bone conduction generator and transmit it to the user.

[0129] Figure 5 FIG1 shows an internal structure diagram of a computer device in an embodiment. The computer device can be a terminal or a server. Figure 5 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement an interactive method based on a visual headset. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement an interactive method based on a visual headset. It will be understood by those skilled in the art that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0130] In one embodiment, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0131] Determining whether a user wearing the visualization headset issues an interaction instruction;

[0132] Analyzing the content of the interactive instructions to determine user needs;

[0133] In response to the user's demand, the camera is controlled to scan the current content to obtain a scanned image;

[0134] Performing data processing on the scanned image based on the user's needs to obtain a processing result;

[0135] Feedback the processing result to the user.

[0136] Through interactive commands, the headset identifies user needs and provides intelligent, personalized services based on them. Through the headset's integrated camera, users can obtain real-time environmental information without relying on external devices and receive instant feedback through the headset, greatly enhancing the practicality and interactive experience of smart wearable devices in complex environments.

[0137] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the processor performs the following steps:

[0138] Determining whether a user wearing the visualization headset issues an interaction instruction;

[0139] Analyzing the content of the interactive instructions to determine user needs;

[0140] In response to the user's demand, the camera is controlled to scan the current content to obtain a scanned image;

[0141] Performing data processing on the scanned image based on the user's needs to obtain a processing result;

[0142] Feedback the processing result to the user.

[0143] Through interactive commands, the headset identifies user needs and provides intelligent, personalized services based on them. Through the headset's integrated camera, users can obtain real-time environmental information without relying on external devices and receive instant feedback through the headset, greatly enhancing the practicality and interactive experience of smart wearable devices in complex environments.

[0144] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0145] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0146] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. An interactive method based on visual headphones, characterized in that: The visualization headset includes a main control center, a headset body, and a camera, wherein the headset body and the camera are respectively connected to the main control center. The method includes: Determining whether a user wearing the visualization headset issues an interaction instruction; Analyzing the content of the interactive instructions to determine user needs; In response to the user's demand, the camera is controlled to scan the current content to obtain a scanned image; Performing data processing on the scanned image based on the user's needs to obtain a processing result; Feedback the processing result to the user.

2. The interactive method based on visual earphones according to claim 1, characterized in that: The visualization headset further includes a light beam transmitter connected to the main control center. The step of performing data processing on the scanned image based on the user's needs to obtain a processing result includes: Determining whether the user has issued a first instruction; If the first instruction is issued, marking the content in the scanned image based on the light beam emitter; determining whether the user has issued a second instruction; If the second instruction is issued, stopping marking the content in the scanned image and extracting the marked content; Data analysis is performed on the marked content based on the user demand to obtain a processing result.

3. The interactive method based on visual earphones according to claim 2, characterized in that: If the second instruction is issued, the step of stopping marking the content in the scanned image and extracting the marked content includes: Acquire a first position in the scanned image where the light beam emitter emits light after the first instruction is issued, and a second position in the scanned image where the light beam emitter emits light after the second instruction is issued; The content from the first position to the second position is marked, and the marked content is extracted.

4. The interactive method based on visual earphones according to claim 1, characterized in that: The step of analyzing the content of the interaction instruction to determine the user's needs includes: Converting the content of the interactive instruction into text information to obtain an interactive text instruction; The interactive text instructions are semantically analyzed through a natural language processing model to obtain user needs.

5. The interactive method based on visual earphones according to claim 1, characterized in that: The step of performing data processing on the scanned image based on the user's needs to obtain a processing result includes: Determining associated feature information based on the user needs; Utilizing a preset computer vision technology, and based on the feature information, extracting key information from the scanned image; The key information is processed based on the user's needs to obtain a processing result.

6. The interactive method based on visual earphones according to claim 1, characterized in that: The feature information includes at least one of text features, picture features, physical features, and QR code features.

7. The interactive method based on visual earphones according to claim 1, characterized in that: The visualization headset further includes a bone conduction generator, and the step of feeding back the processing result to the user includes: sending the processing result to the bone conduction generator; The bone conduction generator converts the processing result into a mechanical signal and transmits it to the user.

8. An interactive device based on a visual headset, characterized in that: The visual headset includes a main control center, a headset body, and a camera. The headset body and the camera are respectively connected to the main control center. The device includes: a judgment module, configured to judge whether the user wearing the visualization headset has issued an interaction instruction; An analysis module, configured to analyze the content of the interaction instruction to determine user needs; A response module, configured to control the camera to scan the current content in response to the user's demand to obtain a scanned image; A processing module, configured to perform data processing on the scanned image based on the user's needs to obtain a processing result; A feedback module is used to feed back the processing result to the user.

9. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is executed by a processor, the processor is caused to perform the steps of the interactive method based on a visual headset according to any one of claims 1 to 7.

10. A computer device, characterized in that: The device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the interactive method based on visual headphones according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Headset intelligent device and interactive system with same

    CN104182051A

  • Intelligent robot control system, method and device based on artificial intelligence

    CN104965426A

  • Intelligent earphone device with computer vision

    CN112073866A

Cited By

  • File travel immersive experience construction system based on digital text creation interaction

    CN120780194A

  • A system for constructing immersive cultural tourism experiences based on digital cultural and creative interactions

    CN120780194B

  • Interaction method and device based on earphone camera, earphone and storage medium

    CN122331858A

  • Earphone camera-based interaction method and apparatus, earphone, and storage medium

    CN122331858B

  • Earphone camera-based interaction method, earphone, and storage medium

    CN122387409A