Intelligent voice interaction device for legal consultation

By introducing multi-layered soundproof curtains, document recognition and projection systems, and emotion perception feedback mechanisms into the legal consultation device, the issues of privacy and user emotion perception in an open environment are resolved, and simplified document interaction and emotion-adaptive feedback are achieved.

CN121932049APending Publication Date: 2026-04-28XIAN JINJU ENTERPRISE MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN JINJU ENTERPRISE MANAGEMENT CO LTD
Filing Date
2025-12-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing legal consultation devices lack the privacy of voice interaction in open environments, the physical document interaction process is cumbersome, and they cannot perceive or respond to changes in the user's emotions.

Method used

The system employs a multi-layered composite soundproof curtain to provide acoustic isolation, and uses a document recognition camera and laser projector to directly overlay digital information with physical documents. It also utilizes facial recognition and voice recognition systems to sense user emotions and provide adaptive feedback.

Benefits of technology

It achieves privacy in voice interaction in an open environment, simplifies the digitization of physical documents, provides adaptive feedback based on user emotions, reduces the risk of information leakage, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121932049A_ABST
    Figure CN121932049A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent interaction devices, and discloses a legal consultation intelligent voice interaction device which comprises an interaction device body, a base, a top seat, a sound insulation mechanism, a pacifying system, a placing glass cover and a scanning projection mechanism. The sound insulation mechanism drives a multi-layer composite sound insulation curtain to be automatically unfolded through a motor, a physically-enclosed private consulting space is formed around a user, sound transmission is blocked, and the privacy of voice information of the user is protected. The scanning projection mechanism obtains a digital image of a physical file through a file recognition camera, and directly projects analyzed digital annotation information to a physical position, corresponding to the image content, on the file through a laser projector, so that superposition display of digital information and a physical entity is achieved, and the display efficiency is improved. The pacifying system collects facial expression and acoustic feature information of a user through a face recognition camera and a voice recognizer, analyzes and judges the emotional state of the user, controls a voice player to play corresponding pacifying voice, and intervenes the emotion of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent interactive device technology, specifically to an intelligent voice interactive device for legal consultation. Background Technology

[0002] Intelligent voice interaction devices, as terminal equipment providing information inquiry and business processing, have been applied in various service scenarios. In professional fields such as legal consultation, these devices can provide users with initial information guidance and document processing support. However, existing technical solutions still have some unresolved technical issues in practical applications.

[0003] When these interactive devices are deployed in public or semi-open environments (such as service halls and community centers), their voice interaction processes typically lack acoustic isolation measures. Sensitive information such as personal identification details and case details mentioned by users during consultations can be directly transmitted through the air, posing a risk of information leakage if overheard by unauthorized personnel. Currently, there is a lack of a structure integrated into the device itself that can generate temporary private spaces on demand.

[0004] Furthermore, when handling interactive tasks involving physical documents (such as contracts and vouchers), existing devices typically employ a scanning display mode. That is, the device first scans the physical document into a digital image, and then presents the image and related analysis results on a separate display screen. This mode requires the user to frequently switch their gaze and attention between the physical document and the device's display screen to establish a correspondence between the two, which increases the cognitive load on the user during understanding and operation.

[0005] Meanwhile, matters such as legal consultations are often accompanied by complex emotional changes in users, such as anxiety or agitation due to cumbersome procedures or difficult-to-understand content. Existing interactive devices merely serve as tools for executing instructions, with fixed interaction logic and lacking the ability to perceive the user's emotional state. Therefore, the devices cannot adaptively adjust their interaction strategies based on changes in the user's emotions, and the entire interaction process lacks effective feedback on the user's psychological state. Summary of the Invention

[0006] The technical problem this invention aims to solve is that existing legal consultation devices are inadequate in providing privacy, handling physical document interaction, and paying attention to the user's emotional state. For example, voice interaction in open environments can easily leak user privacy; the digitization and analysis of physical credentials is cumbersome and cannot achieve intuitive linkage with physical entities; and, in serious legal consultation scenarios, devices generally lack mechanisms for perceiving and alleviating users' negative emotions.

[0007] To address the aforementioned technical problems, this invention provides a legal consultation intelligent voice interaction device, comprising an interaction device body, a base fixedly connected to the bottom of the interaction device body, and a top seat fixedly connected to the top of the interaction device body. A soundproofing mechanism is provided on the outer side of the interaction device body, a soothing system is provided on the outer side of the top seat, and a glass cover is fixedly connected to the outer side of the interaction device body. A scanning projection mechanism is provided on the outer side of the glass cover.

[0008] The sound insulation mechanism includes an unwinding box, which is fixedly connected to the top of the top base. An unwinding rod is rotatably connected inside the top base. Two partitions are fixedly connected to the inner wall of the unwinding box, and two spring-loaded springs are sleeved on the outer side of the unwinding rod. A multi-layer composite sound insulation curtain is fixedly connected to the outer side of the unwinding rod, and two sliders are fixedly connected to the outer side of the multi-layer composite sound insulation curtain. A winding rope is fixedly connected to the outer side of the sliders. Two connecting frames are fixedly connected to the outer side of the unwinding box, and a support guide frame is fixedly connected to the outer side of the connecting frames. A winding box is fixedly connected to the outer side of the connecting frames. A motor is fixedly connected to the inner wall of the winding box, and a winding rod is fixedly connected to the output end of the motor. Two winding wheels are fixedly connected to the outer side of the winding rod.

[0009] In one specific embodiment, the scanning projection mechanism includes a mounting block, the top of which is fixedly connected to the bottom of the support guide frame. A first linear module is disposed at the bottom of the mounting block, a second linear module is disposed at the bottom of the first linear module, and a mounting base is disposed at the bottom of the second linear module. A document recognition camera is mounted at the bottom of the mounting base, and a laser projector is mounted at the top of the mounting base.

[0010] Preferably, the movement directions of the first linear module and the second linear module are perpendicular to each other; the second linear module is disposed on the moving part of the first linear module, and the mounting base is disposed on the moving part of the second linear module. This structure forms a two-dimensional motion platform, enabling the mounting base and the document recognition camera and laser projector it carries to move within the working plane above the placement glass cover.

[0011] Furthermore, in order to accurately map the target region in the document image acquired by the document recognition camera to its physical location on the document, it is necessary to establish a coordinate transformation relationship between the image coordinate system of the document recognition camera and the projection coordinate system of the laser projector. This transformation relationship essentially involves finding a mathematical mapping f such that a point P in the image... cam A point P that can be converted into a target point of the projector proj ,Right now:

[0012] P proj =f(P cam );

[0013] Among them, point P cam Defined by its pixel coordinates (u,v) in the image coordinate system, i.e., P cam (u,v); Point P proj The laser deflection angle (θ) corresponding to it in the projected coordinate system x ,θ y Definition, i.e., P proj (θ x ,θ y In this embodiment, since the image plane and the projection plane are coplanar, the above mapping relationship f can be described by a 3×3 homography transformation matrix H. To perform the calculation using linear algebra, we convert the two-dimensional Cartesian coordinates to three-dimensional homogeneous coordinates. Therefore, the pixel coordinates (u,v) are transformed into projection angle coordinates (θ). x ,θ y The mathematical model of ) is expressed as follows:

[0014]

[0015] in:

[0016] u and v: are the midpoints of the image, respectively. cam The horizontal and vertical pixel coordinates of (u,v).

[0017] It is point P cam The homogeneous coordinate representation of pixel coordinates (u,v);

[0018] θ x and θ y : These are the points P in the image. cam The physical location corresponds to the target angle coordinates required for the laser projector micromirror system to deflect; these two coordinates together define the projection point P. proj (θ x ,θ y );

[0019] w: is a non-zero scale factor generated in homogeneous coordinate operations, used for subsequent normalization calculations;

[0020] It is point P proj Target angle coordinates (θ) x ,θ y Homogeneous coordinate representation of ).

[0021] H: is the 3×3 homography transformation moment describing the mapping relationship from the image coordinate system to the projected coordinate system;

[0022] h 11 ,h 12 ,...,h32 : are the eight independent parameters of matrix H, whose values ​​are obtained by collecting multiple pairs of pixel coordinates and projection angle coordinates during the calibration process and calculating them using algorithms such as direct linear transformation.

[0023] In practice, the document recognition camera captures images of documents to obtain document images containing the target region. The interactive device analyzes the document image to determine the pixel coordinates (u, v) of the target region. Subsequently, using a pre-calibrated matrix H, the corresponding projection angle coordinates (θ) are calculated. x ,θ y Finally, the main body of the interactive device controls the micromirror system of the laser projector to deflect at the calculated angle, accurately projecting the annotation information onto the physical document at the position corresponding to the target area.

[0024] In one specific embodiment, the multi-layer composite soundproof curtain consists of, from the outside to the inside, a decorative fabric made of polyester fiber, a high-density soundproof felt made of polymer, and a skin-friendly soft material made of silicone.

[0025] In one specific embodiment, the soothing system includes a facial recognition camera disposed on the outer side of the top mount, a voice recognizer disposed inside the main body of the interactive device, and a voice player. The facial recognition camera is used to collect the user's facial expression information, and the voice recognizer is used to collect the user's voice information. The main body of the interactive device analyzes and judges the user's emotional state based on the facial expression information and the voice information, and controls the voice player to play soothing voice messages corresponding to the emotional state.

[0026] Preferably, when the face recognition camera recognizes the user and the voice recognition device receives a preset activation command, the main body of the interactive device controls the motor to rotate to unfold the multi-layer composite soundproof curtain.

[0027] In one specific embodiment, the end of the take-up rope away from the slider is fixedly connected to the outside of the take-up reel, and the take-up rod is rotatably connected inside the take-up box.

[0028] In one specific embodiment, one end of the spring is fixedly connected to the outside of the partition, and the other end of the spring is fixedly connected to the inner wall of the unwinding box.

[0029] In one specific embodiment, the connecting frame has a groove inside, and the slider is slidably connected to the inner wall of the groove.

[0030] This invention provides an intelligent voice interaction device for legal consultation. It has the following beneficial effects:

[0031] 1. This invention utilizes a soundproofing mechanism, employing a motor to automatically unfold a multi-layered composite soundproof curtain, creating a physically enclosed space around the user. Because the multi-layered composite soundproof curtain contains materials such as high-density soundproof felt, it effectively attenuates the propagation intensity of voice signals in the air, thus providing users with a temporary, acoustically isolated consultation area in an open environment, reducing the risk of user voice information being leaked to the outside.

[0032] 2. This invention utilizes a scanning projection mechanism and a document recognition camera to acquire digital images of physical documents. These images are then analyzed by the interactive device to locate the target area. Subsequently, a laser projector, controlled by a pre-established coordinate transformation relationship, directly projects the digital annotation information onto the physical document at the physical location corresponding to the target area. This structure enables the direct overlay display of digital analysis results and physical documents, eliminating the need for the user to switch their gaze between the device screen and the physical document.

[0033] 3. This invention employs a soothing system that utilizes a facial recognition camera and a voice recognizer to collect the user's facial expressions and acoustic features. The main body of the interactive device fuses and analyzes this multimodal information to determine the user's current emotional state, and based on the determination result, controls the voice player to play preset soothing voice messages corresponding to that emotional state. This device possesses the ability to perceive and respond to the user's emotional state. Attached Figure Description

[0034] Figure 1 This is a perspective view of the present invention;

[0035] Figure 2 This is a schematic diagram of the voice player of the present invention;

[0036] Figure 3 This is a schematic diagram of the multi-layer composite soundproof curtain of the present invention;

[0037] Figure 4 This is a schematic diagram of the unwinding box of the present invention;

[0038] Figure 5 This is a schematic diagram of the winding box of the present invention;

[0039] Figure 6 for Figure 5 Enlarged view of point B;

[0040] Figure 7 for Figure 5 Enlarged view of point A;

[0041] Figure 8 This is a schematic diagram of the laser projector of the present invention;

[0042] Figure 9 This is a schematic diagram of the high-density sound insulation felt of the present invention.

[0043] The components include: 1. Interactive device main body; 2. Base; 3. Top seat; 4. Support guide frame; 5. Connecting frame; 6. Unwind box; 7. Partition; 8. Spring; 9. Unwind rod; 10. Multi-layer composite soundproof curtain; 11. Slider; 12. Winding rope; 13. Winding box; 14. Winding rod; 15. Winding wheel; 16. Motor; 17. Voice player; 18. Placement glass cover; 19. Face recognition camera; 20. Mounting block; 21. First linear module; 22. Second linear module; 23. Mounting base; 24. Document recognition camera; 25. Laser projector; 26. Decorative fabric; 27. High-density soundproof felt; 28. Skin-friendly soft material; 29. ​​Voice recognizer; 30. Slide. Detailed Implementation

[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] Please see the appendix Figure 1 -Appendix Figure 9 This invention provides a legal consultation intelligent voice interaction device, which physically includes an interaction device body 1. The bottom of the interaction device body 1 is fixedly connected to a base 2 to ensure the stability of the device. A top seat 3 is fixedly connected to the top of the interaction device body 1.

[0046] The main body 1 of the interactive device is a hollow shell structure, which houses the central control and processing unit of the device. In this embodiment, the central control and processing unit is defined as a processor module. The processor module includes at least one or more processors, a memory for storing program instructions and data, and an input / output interface for signal transmission with other components. A glass cover 18 is fixedly connected to one outer wall of the main body 1 of the interactive device, which provides a flat, transparent surface for the user to place physical documents.

[0047] The three core functional mechanisms in this embodiment—the sound insulation mechanism, the soothing system, and the scanning projection mechanism—are all arranged around the main body 1 of the interactive device. Specifically, the main components of the sound insulation mechanism, including the unwinding box 6, the rewinding box 13, and the connecting frame 5 connecting them, are all located on the top of the top seat 3. The sensor components of the soothing system, such as the face recognition camera 19, are also located on the outside of the top seat 3, while its execution components, such as the voice player 17, and signal acquisition components, such as the voice recognizer 29, are located inside the main body 1 of the interactive device.

[0048] The scanning projection mechanism is located on the outside of the glass cover 18. Specifically, it is suspended by a support guide 4. The support guide 4 is fixedly connected to the outside of the connecting frame 5, thereby structurally linking the scanning projection mechanism to the top base 3.

[0049] To achieve coordinated control of the various functional mechanisms, a processor module housed within the main body 1 of the interactive device establishes electrical connections with the electronic components of each functional mechanism. Specifically, the processor module is electrically connected to the motor 16 of the soundproofing mechanism, the first linear module 21 and the second linear module 22 of the scanning projection mechanism, the document recognition camera 24 and the laser projector 25, and the face recognition camera 19, voice recognition device 29, and voice player 17 of the soothing system. Through these connections, the processor module can send control commands and receive data to execute all the functions of the present invention.

[0050] Please see the appendix Figure 1 -Appendix Figure 5 The various components of the sound insulation mechanism work together to achieve acoustic isolation around the user consultation area. The fixed part of the sound insulation mechanism is fixedly connected to the top of the top seat 3 by two connecting brackets 5. The two connecting brackets 5 are arranged parallel to each other, and each has a sliding groove 30 on its opposite inner sidewall.

[0051] A roll-up box 6 is fixedly installed at one end of each of the two connecting frames 5. A roll-up rod 9 is rotatably connected inside the roll-up box 6. One end of the multi-layer composite soundproof curtain 10 is fixedly connected to the roll-up rod 9 and is stored in a rolled-up state. Two partitions 7 are fixedly connected to the inner wall of the roll-up box 6. One end of a spring 8 is fixed to the outside of the partition 7, and the other end is fixed to the roll-up rod 9. The spring 8 is passively wound up to store potential energy when the multi-layer composite soundproof curtain 10 is unfolded.

[0052] At the other end of the two connecting brackets 5, a take-up box 13 is fixedly installed. A motor 16 is fixedly installed inside the take-up box 13, and a take-up rod 14 is fixedly connected to the output end of the motor 16. Two take-up wheels 15 are fixedly sleeved on the take-up rod 14.

[0053] The free end of the multi-layer composite soundproof curtain 10 is fixedly connected between two sliders 11. The sliders 11 are slidably connected inside the grooves 30 of the two connecting frames 5, so that the free end of the multi-layer composite soundproof curtain 10 can only move linearly along the trajectory of the grooves 30. One end of the winding rope 12 is fixedly connected to the slider 11, and the other end passes around the winding wheel 15 and is fixed.

[0054] When the multi-layer composite soundproof curtain 10 needs to be unfolded, the processor module located inside the main body 1 of the interactive device sends a forward rotation command to the motor 16. The motor 16 drives the take-up rod 14 and the take-up wheel 15 to rotate, thereby winding and tightening the take-up rope 12. The take-up rope 12 generates a pulling force on the slider 11, driving the slider 11 to move along the slide groove 30 from one end of the unwinding box 6 to the other end of the take-up box 13. During this process, the slider 11 pulls the multi-layer composite soundproof curtain 10 to unfold from the unwinding rod 9, while simultaneously forcing the unwinding rod 9 to rotate, causing the spring 8 to be passively tightened.

[0055] When the multi-layer composite soundproof curtain 10 needs to be wound up, the processor module sends a reverse rotation command to the motor 16 or puts it into a free state. The potential energy stored in the spring 8 is released, generating a restoring torque that acts on the unwinding rod 9. This torque drives the unwinding rod 9 to rotate in the reverse direction, rewinding the multi-layer composite soundproof curtain 10 onto it. The tension generated during the winding process acts on the slider 11, causing it to return to its original position along the groove 30 towards the unwinding box 6. During this process, the slider 11 drives the winding rope 12, causing the winding wheel 15 and the winding rod 14 to rotate in the opposite direction, thereby releasing the winding rope 12.

[0056] In this embodiment, the multi-layer composite soundproof curtain 10 is a layered composite structure, in which decorative fabric 26, high-density soundproof felt 27, and skin-friendly soft material 28 are sequentially stacked from the outside to the inside. The decorative fabric 26 is made of polyester fiber material, providing structural strength and aesthetic appeal. The high-density soundproof felt 27 is made of polymer material, and its high density and internal damping properties are used to absorb and block sound wave energy. The skin-friendly soft material 28 is made of silicone material, providing a soft inner surface to prevent discomfort from accidental contact by the user.

[0057] Please see the appendix Figure 1 and attached Figure 8 The scanning projection mechanism enables non-contact scanning and information projection of physical documents placed on the glass cover 18. This mechanism is suspended from the bottom of the support guide frame 4 by a mounting block 20, which is fixedly connected to the outside of the connecting frame 5.

[0058] The bottom of the mounting block 20 is fixedly connected to the stationary base of the first linear module 21. The first linear module 21 includes a guide rail extending along a first direction and a moving component that can move linearly on the guide rail. The stationary base of the second linear module 22 is fixedly mounted on the moving component of the first linear module 21.

[0059] The second linear module 22 also includes a guide rail and a moving part that can move linearly thereon. The extension direction of the guide rail of the second linear module 22 is perpendicular to the extension direction of the guide rail of the first linear module 21. The mounting base 23 is fixedly mounted on the moving part of the second linear module 22. Through this orthogonal stacking arrangement, the first linear module 21 and the second linear module 22 together constitute a two-dimensional motion platform.

[0060] Mounting base 23 is a shared support component for both the document recognition camera 24 and the laser projector 25. Specifically, the document recognition camera 24 is mounted at the bottom of mounting base 23, with its lens optical axis pointing vertically downwards and aligned with the glass cover 18 below. The laser projector 25 is mounted at the top of mounting base 23, with its laser projection optical axis also pointing vertically downwards. Due to their different physical mounting positions, the center of the lens of the document recognition camera 24 and the projection center of the laser projector 25 have a fixed offset in the horizontal plane.

[0061] During operation, the processor module located inside the main body 1 of the interactive device sends control signals to the drive motors inside the first linear module 21 and the second linear module 22, respectively. When the processor module controls the first linear module 21 to work, its moving parts drive the entire second linear module 22 and the mounting base 23 fixed on it to move along the first direction. When the processor module controls the second linear module 22 to work, its moving parts only drive the mounting base 23 to move along the second direction perpendicular to the first direction. By coordinating the control of the two linear modules, the processor module can drive the mounting base 23 to any specified coordinate position within the working plane above the glass cover 18, thereby achieving precise positioning of the document recognition camera 24 and the laser projector 25.

[0062] Since the document recognition camera 24 and the laser projector 25 are physically two independent components, their optical centers are offset in their physical installation. Therefore, there is an inherent transformation relationship between the image coordinate system of the document recognition camera 24 and the projection coordinate system of the laser projector 25. To accurately map the target area in the image acquired by the document recognition camera 24 onto the physical document via the laser projector 25, a precise mathematical mapping between these two coordinate systems must be established. The goal of this mathematical mapping is to map any point P in the image... cam Transformed into a target point P of the projector proj In this embodiment, point P cam It is uniquely defined by its pixel coordinates (u,v) in the image, denoted as P. c am(u,v); and point P proj The horizontal and vertical angles (θ) required for the internal micromirror system of the driving laser projector 25 to deflect are then determined. x ,θ yLet P be a unique definition. proj )θ x ,θ y Because the image acquisition plane and the laser projection plane are both on the same physical plane where the glass cover 18 is placed, point P... cam With point P proj The transformation between them is a projective transformation from a two-dimensional plane to another two-dimensional plane, and this transformation relationship can be described by a 3×3 homography transformation matrix H. To utilize linear algebra for calculation, the two-dimensional coordinates are converted to three-dimensional homogeneous coordinates, and the mathematical model is as follows:

[0063]

[0064] The parameters in the above mathematical model are defined as follows:

[0065] The input vector originates from point P. cam ):

[0066] u: Midpoint P in the image c The horizontal pixel coordinate of am.

[0067] v: Midpoint P in the image c The vertical pixel coordinates of am.

[0068] From point P c The homogeneous coordinate vector formed by the pixel coordinates (u,v) of am;

[0069] From point P c The homogeneous coordinate vector formed by the pixel coordinates (u,v) of am

[0070] Transformation matrix:

[0071] H is a 3×3 homography transformation matrix whose internal parameters define the complete mapping relationship from the image coordinate system to the projected coordinate system;

[0072] h 11 ,h 12 ,...,h 32 These are the eight independent parameters of the homography transformation matrix H, and their specific values ​​are determined by subsequent calibration procedures.

[0073] Output vector (used to define point P) proj ):

[0074] θ x : and point P p The horizontal deflection angle corresponding to roj;

[0075] θ y : and point P pThe vertical deflection angle corresponding to roj;

[0076] w: A non-zero scale factor generated during homogeneous coordinate calculation;

[0077] From point P p roj's target angle coordinates (θ) x ,θ y The homogeneous coordinate vector formed by ).

[0078] The specific parameters of the homography transformation matrix H are determined through a calibration process. This process is executed by a processor module located inside the main body 1 of the interactive device, and the specific steps are as follows: First, the processor module controls the laser projector 25 to move according to a set of preset deflection angle coordinates {(θ... x,i ,θ y,i A series of light spots are sequentially projected onto the glass cover 18. Then, for each projected light spot, the processor module controls the file recognition camera 24 to capture one frame of image. The processor module then performs image processing algorithms on the image, such as image binarization and centroid calculation, to accurately obtain the pixel coordinates {(u} of the light spot in the image. i ,v i By repeating this process, the processor module obtains at least four sets of paired points that correspond one-to-one with the projection angle coordinates and image pixel coordinates.

[0079] Finally, the processor module substitutes these paired point sets into the mathematical model of homography transformation, forming an overdetermined system of linear equations. It then uses a direct linear transformation algorithm to solve this system of equations, thereby calculating the eight parameter values ​​of matrix H. The calculated matrix H is stored in the processor module's non-volatile memory.

[0080] In the actual operation of the device, when it is necessary to select a target point (u) in the image target ,v target When annotating the physical file, the processor module first reads the pre-calibrated homography transformation matrix H from memory. Then, the processor module performs matrix-vector multiplication to calculate the homogeneous coordinates of the target point in the projected coordinate system.

[0081]

[0082] Next, by dividing by the scale factor w′ to perform dehomogenization, the final laser deflection angle coordinates are obtained:

[0083] θ x =x' / w';

[0084] θ y =y' / w';

[0085] The parameters in the above calculation process are defined as follows:

[0086] u target : The horizontal pixel coordinates of the target point to be labeled in the image;

[0087] u target : The vertical pixel coordinates of the target point to be labeled in the image;

[0088] byu target and u target The intermediate homogeneous coordinate vector obtained by matrix H transformation;

[0089] x′: The first element of the intermediate vector, i.e., h 11 u target +h 12 v target +h 13 .

[0090] y′: The second element of the intermediate vector, i.e., h 21 u target +h 22 v target +h 23 ;

[0091] w′: The third element of the intermediate vector, i.e., h 31 u target +h 32 v target +1 is the scale factor obtained from this specific calculation;

[0092] x′ / w′: Dehomogenization operation, dividing the first element of the intermediate vector by the third element to obtain the final horizontal deflection angle θ. x ;

[0093] y′ / w′: Dehomogenization operation, dividing the second element of the intermediate vector by the third element to obtain the final vertical deflection angle θ. y .

[0094] Finally, the processor module will calculate (θ) x ,θ y The laser beam is sent as an instruction to the drive circuit of the laser projector 25, which controls its internal micromirror system to precisely deflect to the target angle, so that the laser beam is accurately projected onto the physical document and the target point (u) of the image. target ,v target The corresponding physical location.

[0095] Please see the appendix Figure 1 and attached Figure 2The soothing system is used to sense the user's emotional state and provide voice feedback. Physically, the system includes a face recognition camera 19 located on the outside of the top mount 3, a voice recognizer 29 located inside the main body 1 of the interactive device, and a voice player 17. All three components are electrically connected to a processor module located inside the main body 1 of the interactive device.

[0096] The information flow begins with data acquisition. A face recognition camera 19 is pointed towards the user's area to continuously or at preset time intervals capture digital image frames containing the user's face. Simultaneously, a voice recognizer 29 uses its built-in microphone to collect the user's voice during interaction and converts the sound wave signal into a digital audio signal stream. These two independent digital signal streams are transmitted to the processor module in real time.

[0097] The second stage of the information flow is feature extraction. After receiving a digital image frame, the processor module executes a face detection algorithm to locate facial regions in the image and further extracts a set of predefined facial keypoint coordinates within that region. These keypoints describe the geometry and relative positions of the eyebrows, eyes, nose, and mouth. For the received digital audio signal stream, the processor module executes an acoustic analysis algorithm to extract a set of acoustic feature vectors, which contain parameters such as fundamental frequency, energy, speech rate, and Mel-frequency cepstral coefficients.

[0098] The third stage of the information flow is state determination. The processor module combines the extracted facial key point coordinate features and acoustic feature vectors to form a multimodal feature vector. An emotion state classification model is pre-stored in the processor module's memory. The multimodal feature vector is used as input to this model. The model calculates the input vector, and its output is a classification result representing a specific emotion state, such as classifying the result into a preset discrete category like "calm," "anxious," or "excited."

[0099] The final stage of the information flow is the execution feedback. Based on the obtained emotional state classification results, the processor module retrieves a soothing voice file corresponding to the state from a pre-defined lookup table or database. Once the corresponding voice file is found, the processor module sends the file's digital audio data along with a playback command to the voice player 17 via an electrical connection. Upon receiving the command and data, the voice player 17 drives the speaker through its digital-to-analog converter and amplification circuit, converting the digital audio data into sound for playback.

[0100] The emotion recognition and feedback principle of the soothing system is achieved through a series of algorithms in the processor module to analyze and make decisions on the collected multimodal data. Its principle includes steps such as visual feature extraction, acoustic feature extraction, multimodal feature fusion, emotion state classification, and feedback generation.

[0101] In the visual feature extraction step, the processor module processes image frames containing the user's face acquired from the face recognition camera 19. First, a face detection algorithm is executed to determine the bounding box of the face in the image. Then, within the located facial region, a facial keypoint detection algorithm is executed to locate a set of predefined, anatomically significant points, such as eyebrow contour points, corners of the eyes, tip of the nose, and corners of the mouth. Based on the coordinates of these keypoints, the processor module calculates a set of geometric features, such as Euclidean distances and angles between the keypoints, as well as the activation intensity of specific action units defined according to the facial motion coding system. These calculated values ​​are combined into a visual feature vector V. face .

[0102] In the acoustic feature extraction step, the processor module processes the digital audio signal acquired from the speech recognizer 29. This signal is segmented into a series of short frames. For each frame, the processor module calculates a set of acoustic features, primarily including: prosodic features such as fundamental frequency, energy, or intensity; and spectral features such as Mel-frequency cepstral coefficients and their first and second-order differences. Then, the frame-level features over a time period (e.g., a complete speech segment) are statistically analyzed, calculating their mean, standard deviation, maximum, minimum, and other statistical measures. These statistical measures are combined into an acoustic feature vector V. audiω In the multimodal feature fusion and emotion state classification steps, the processor module will process the visual feature vector V generated in the first two steps. face Harmony acoustic eigenvector V audio Combine them to form a higher-dimensional fusion feature vector V fusion .

[0103] In one specific implementation, this fusion is accomplished through feature layer concatenation, i.e., V fusion =[V face ⊕V audiω ], where ⊕ represents the vector concatenation operation. Subsequently, the fused feature vector V fusion The input is fed into a pre-trained emotion state classification model. This model can be a support vector machine, a random forest model, or a deep neural network. The model, based on its internal parameters, classifies the input V... fusion The system performs calculations and outputs probability scores for one or more predefined emotion categories (e.g., calm, anxious, agitated). The processor module selects the category with the highest probability score as the final emotion state judgment result C. emotion This process can be represented by the following formula:

[0104]

[0105] in:

[0106] Cemotion The final emotional state class calculated;

[0107] k: Represents a predefined sentiment category index;

[0108] A mathematical operator that takes the argument that maximizes the value of the subsequent expression;

[0109] M represents a pre-trained emotion state classification model, which is a function whose input is a fused feature vector and whose output is a probability score vector for each emotion category.

[0110] V fusion A fused feature vector composed of visual and acoustic features;

[0111] M(V fusion ) k : indicates that the classification model M is given by the input vector V fusion The calculated probability score for belonging to the k-th emotion category.

[0112] In the feedback generation step, the processor module uses the calculated emotion state category C. emotion The query key is used to look up information in an internally stored mapping table. This mapping table predefines the correspondence between each emotional state category and one or more specific soothing voice files. When a match is found with C... emotion After receiving the corresponding voice file, the processor module reads the audio data of the file from the memory and sends it along with a playback command to the voice player, thereby achieving adaptive voice feedback based on the user's current emotional state.

[0113] Working principle: In the initial standby state, the legal consultation intelligent voice interaction device is powered on, but its functional mechanisms are inactive, and the multi-layer composite soundproof curtain 10 is completely rolled up inside the unwinding box 6. When a user approaches the device, the facial recognition camera 19 in the soothing system captures the user's facial image and confirms the user's presence.

[0114] Subsequently, the user issues a preset activation command via voice, which is collected by the voice recognition unit 29 and transmitted to the processor module. After recognizing the activation command, the processor module initiates the device's interactive process.

[0115] The first step in the process is to create a private consultation space. The processor module sends a forward rotation command to the motor 16 of the soundproofing mechanism. The motor 16 drives the winding rod 14 and winding wheel 15 to rotate, tightening the winding rope 12. The tension of the winding rope 12 acts on the slider 11, causing it to move along the groove 30 inside the connecting frame 5, thereby smoothly unfolding the multi-layer composite soundproofing curtain 10 from the unwinding rod 9 until it is completely covered.

[0116] The second step of the process is document scanning and information interaction. When the user places a physical document on the glass cover 18 and issues a scan command, the processor module controls the scanning projection mechanism. The processor module first controls the movement of the first linear module 21 and the second linear module 22 to drive the document recognition camera 24 to scan the entire document or locate a specific area and capture an image. The processor module analyzes the acquired image, for example, locating the target area that the user needs to focus on through optical character recognition, and obtaining the pixel coordinates (u,v) of that area in the image coordinate system. Subsequently, the processor module reads the pre-calibrated homography transformation matrix H from its memory and transforms the pixel coordinates (u,v) into the projection angle coordinates (θ) of the laser projector 25 according to this matrix. x ,θ y Finally, the processor module sends the angular coordinates as an instruction to the laser projector 25, which then projects the preset annotation information, such as highlighted outlines or text annotations, precisely onto the physical document at the physical location that corresponds exactly to the target area.

[0117] Throughout the interaction process, the soothing system runs continuously in the background. The facial recognition camera 19 and voice recognition device 29 periodically collect the user's facial expression data and voice acoustic data. The processor module analyzes this multimodal data in real time to determine the user's current emotional state. When the determination result is a preset negative emotional state such as anxiety or agitation, the processor module automatically retrieves and selects the corresponding soothing voice file from the memory and controls the voice player 17 to play it.

[0118] When the consultation ends and the user issues an end command, the processor module sends a reverse or trip command to the motor 16 of the soundproofing mechanism. At this time, the spring 8, which was passively tightened due to the unfolding of the soundproofing curtain, releases its stored elastic potential energy, generating a restoring torque that drives the unwinding lever 9 to rotate in the opposite direction. This rotational force rewinds the multi-layer composite soundproofing curtain 10 back into the unwinding box 6. During this process, the slider 11 is pulled by the curtain to reset towards one end of the unwinding box 6, and the winding rope 12 is released. All mechanisms of the device return to the initial standby state.

Claims

1. A legal consultation intelligent voice interaction device, characterized in that, The device includes an interactive device body (1), a base (2) is fixedly connected to the bottom of the interactive device body (1), a top seat (3) is fixedly connected to the top of the interactive device body (1), a sound insulation mechanism is provided on the outside of the interactive device body (1), a soothing system is provided on the outside of the top seat (3), a glass cover (18) is fixedly connected to the outside of the interactive device body (1), and a scanning projection mechanism is provided on the outside of the glass cover (18). The sound insulation mechanism includes a roll-up box (6), which is fixedly connected to the top of the top seat (3). A roll-up rod (9) is rotatably connected inside the top seat (3). Two partitions (7) are fixedly connected to the inner wall of the roll-up box (6). Two spring-loaded springs (8) are sleeved on the outside of the roll-up rod (9). A multi-layer composite sound insulation curtain (10) is fixedly connected to the outside of the roll-up rod (9). Two sliders (11) are fixedly connected to the outside of the multi-layer composite sound insulation curtain (10). A winding rope (12) is fixedly connected to the outside of the slider (11). Two connecting frames (5) are fixedly connected to the outside of the unwinding box (6). A support guide frame (4) is fixedly connected to the outside of the connecting frame (5). A winding box (13) is fixedly connected to the outside of the connecting frame (5). A motor (16) is fixedly connected to the inner wall of the winding box (13). A winding rod (14) is fixedly connected to the output end of the motor (16). Two winding wheels (15) are fixedly connected to the outside of the winding rod (14).

2. The intelligent voice interaction device for legal consultation according to claim 1, characterized in that, The scanning projection mechanism includes a mounting block (20), the top of which is fixedly connected to the bottom of the support guide frame (4). A first linear module (21) is provided at the bottom of the mounting block (20), a second linear module (22) is provided at the bottom of the first linear module (21), a mounting base (23) is provided at the bottom of the second linear module (22), a document recognition camera (24) is installed at the bottom of the mounting base (23), and a laser projector (25) is installed at the top of the mounting base (23).

3. The intelligent voice interaction device for legal consultation according to claim 1, characterized in that, The multi-layer composite soundproof curtain (10) consists of, from the outside to the inside, a decorative fabric (26) made of polyester fiber, a high-density soundproof felt (27) made of polymer, and a skin-friendly soft material (28) made of silicone.

4. The intelligent voice interaction device for legal consultation according to claim 1, characterized in that, The reassurance system includes: A face recognition camera (19) is installed on the outside of the top seat (3), a voice recognition device (29) is installed inside the main body (1) of the interactive device, and a voice player (17) is installed. The face recognition camera (19) is used to collect the user's facial expression information, and the voice recognition device (29) is used to collect the user's voice information; The main body (1) of the interactive device analyzes and judges the user's emotional state based on the facial expression information and the voice information, and controls the voice player (17) to play a soothing voice corresponding to the emotional state.

5. The intelligent voice interaction device for legal consultation according to claim 4, characterized in that, When the face recognition camera (19) recognizes the user and the voice recognition device (29) receives the preset activation command, the main body (1) of the interactive device controls the motor (16) to rotate to unfold the multi-layer composite soundproof curtain (10).

6. The intelligent voice interaction device for legal consultation according to claim 2, characterized in that, The movement directions of the first linear module (21) and the second linear module (22) are perpendicular to each other; the second linear module (22) is disposed on the moving part of the first linear module (21), and the mounting base (23) is disposed on the moving part of the second linear module (22).

7. A legal consultation intelligent voice interaction device according to claim 2, characterized in that, The document recognition camera (24) is used to capture the document placed on the glass cover (18) to obtain the document image. The interactive device main body (1) recognizes the document image and locates the target area in the document image according to the recognition result. The interactive device main body (1) controls the laser projector (25) to project the annotation information onto the document at the physical position corresponding to the target area.

8. The intelligent voice interaction device for legal consultation according to claim 1, characterized in that, The end of the take-up rope (12) away from the slider (11) is fixedly connected to the outside of the take-up wheel (15), and the take-up rod (14) is rotatably connected to the inside of the take-up box (13).

9. A legal consultation intelligent voice interaction device according to claim 1, characterized in that, One end of the spring (8) is fixedly connected to the outside of the partition (7), and the other end of the spring (8) is fixedly connected to the inner wall of the unwinding box (6).

10. A legal consultation intelligent voice interaction device according to claim 1, characterized in that, The connecting frame (5) has a groove (30) inside, and the slider (11) is slidably connected to the inner wall of the groove (30).