Voice translation method and system for overall consultation picture
By remote interaction between consultation spaces, the overall consultation picture is constructed, and combined with the use of dynamic matching and recognition learning models, the problem of single-dimensional photography of consultation space in the existing technology is solved, and efficient speech translation effect is achieved.
Patent Information
- Application Number
- CN202510252366.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, shooting of consultation space is limited to a single dimension, resulting in insufficient matching between Chinese information and each segment in the overall consultation picture, and the inability to achieve efficient speech translation.
By remote interaction between the first consultation space of the ambulance and the second consultation space of the hospital, the interaction time period is defined, and the overall consultation picture is constructed based on the pictures taken in this time period and the two spaces. Then, multiple sub-consultation segments are divided, the posture picture and voice segments are dynamically matched, and speech translation is performed based on the person's movements and mouth shape, and abnormal nodes are identified through the recognition learning model for local optimization.
It realizes multi-dimensional control of the overall consultation screen, ensures the high matching between Chinese information and each clip in the consultation screen, and improves the accuracy and effect of speech translation.
Smart Images

Figure CN120091171A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech translation for the overall consultation picture, and in particular to a method and system for speech translation of the overall consultation picture. Background Art
[0002] With the development of technology, for the control of the consultation space and real-time shooting of the consultation space, in the prior art, the consultation space is photographed in a single dimension to form a corresponding consultation picture. However, the consultation picture needs to be manually dubbed later, and the matching of Chinese information and each segment in the overall consultation picture cannot be guaranteed. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art. The present invention provides a method and system for speech translation of the overall consultation picture. The rescue personnel conduct a primary consultation on the wounded in the first consultation space of the ambulance and collect the preliminary results of the primary consultation; remotely interact with the first consultation space of the ambulance and the second consultation space of the hospital, and define the secondary consultation results according to the medical staff in the second consultation space and the preliminary results of the primary consultation; define the interaction time period based on the time point where the preliminary results of the primary consultation are located and the time point where the secondary consultation results are located; construct the overall consultation picture according to the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, which comprehensively considers the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, and multi-dimensionally controls the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, ensuring the accuracy of the overall consultation picture.
[0004] Furthermore, divide the overall consultation picture into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the gesture picture and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person; match the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, define the abnormal nodes according to the recognition learning model and the recognition of the overall consultation picture, and perform local optimization on the abnormal nodes to improve the overall consultation picture. Multiple mechanisms are carried out for the output of the overall consultation picture and the recognition of the overall consultation picture, so as to realize multi-level control of the overall consultation picture, ensure the high matching of Chinese information and each segment in the overall consultation picture, and improve the speech translation effect of the overall consultation picture.
[0005] An embodiment of the present invention provides a method for speech translation of the overall consultation picture, which is applied to the speech translation scenario of the overall consultation picture;
[0006] The method for voice translation of the overall consultation picture includes:
[0007] The rescue personnel conduct a primary consultation on the wounded based on the first consultation space of the ambulance and collect the preliminary results of the primary consultation.
[0008] The first consultation space of the remote interaction ambulance and the second consultation space of the hospital are interacted, and the secondary consultation results are defined according to the medical staff in the second consultation space and the preliminary results of the primary consultation.
[0009] An interaction time period is defined based on the time point where the preliminary results of the primary consultation are located and the time point where the secondary consultation results are located.
[0010] An overall consultation picture is constructed according to the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space.
[0011] Based on the overall consultation picture, multiple sub-consultation segments are divided, and the Chinese information of each sub-consultation segment is output based on each sub-consultation segment. At this time, the posture picture and the voice segment are dynamically matched, and the Chinese information described by each person is translated based on the association of the actions, lip shapes, and corresponding voices of each person.
[0012] According to the overall consultation picture, the Chinese information, and the action time nodes of each person, the corresponding recognition learning model is matched. According to the recognition learning model and the recognition of the overall consultation picture, abnormal nodes are defined, and local optimization is performed on these abnormal nodes to improve the overall consultation picture.
[0013] Optionally, the rescue personnel conduct a primary consultation on the wounded based on the first consultation space of the ambulance and collect the preliminary results of the primary consultation, including:
[0014] Collect the arrival signal of the ambulance, locate the rescue position according to the analysis of the arrival signal of the ambulance, and collect the accident segment according to the detection of the rescue position.
[0015] Define the accident type based on the accident segment and the accident learning model, and define the preliminary injury of the wounded according to the accident type, the injured segment of the wounded, and the duration. Transmit the preliminary injury of the wounded to the first consultation space of the ambulance.
[0016] Define the injury form according to the preliminary injury of the wounded and the actual injured image of the wounded, and perform preliminary treatment according to the injury form and the rescue tools available in the first consultation space. The rescue personnel conduct a primary consultation on the wounded based on the first consultation space of the ambulance.
[0017] Collect the preliminary results of the primary consultation.
[0018] Optionally, the first consultation space of the remote interaction ambulance and the second consultation space of the hospital define the secondary consultation result according to the medical staff in the second consultation space and the preliminary result of the primary consultation, including:
[0019] Freeze the first consultation space of the ambulance;
[0020] Collect the second consultation space of the hospital;
[0021] Conduct remote interaction for the second consultation space of the hospital and the first consultation space of the ambulance;
[0022] Freeze the preliminary result of the primary consultation;
[0023] During the remote interaction, the medical staff in the second consultation space make a preliminary judgment on the preliminary result of the primary consultation and output the corresponding judgment result;
[0024] Associate the judgment result, the preliminary result of the primary consultation, and the injury form of the wounded;
[0025] Conduct multiple interactions on the judgment result, the preliminary result of the primary consultation, and the injury form of the wounded, and define the corresponding secondary consultation result based on the judgment result, the preliminary result of the primary consultation, and the injury form of the wounded.
[0026] Optionally, defining the interaction time period based on the time point where the preliminary result of the primary consultation is located and the time point where the secondary consultation result is located, including:
[0027] Freeze the preliminary result of the primary consultation and the secondary consultation result;
[0028] Collect the time point where the preliminary result of the primary consultation is located, and at the same time, collect the time point where the secondary consultation result is located;
[0029] Associate the time point where the preliminary result of the primary consultation is located and the time point where the secondary consultation result is located;
[0030] Define the interaction time period according to the time point where the preliminary result of the primary consultation is located and the time point where the secondary consultation result is located;
[0031] Define the abnormal time based on the traversal of the interaction time period, define the corresponding abnormal picture based on the trace of the abnormal time, and optimize the abnormal time according to the intelligent recognition of the abnormal picture to improve the interaction time period.
[0032] Optionally, constructing the overall consultation picture according to the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space, including:
[0033] Freeze the interaction time period;
[0034] Associate the interaction time period, the images captured in the first consultation space, and the images captured in the second consultation space;
[0035] Define a first sub-image based on the interaction time period and the images captured in the first consultation space;
[0036] Define a second sub-image based on the interaction time period and the images captured in the second consultation space, and define a transition image based on the image matching between the second sub-image and the first sub-image;
[0037] Construct an overall consultation image based on the first sub-image, the second sub-image, and the transition image.
[0038] Optionally, divide the overall consultation image into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture image and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person, including:
[0039] Freeze the overall consultation image;
[0040] Divide the overall consultation image into multiple sub-consultation segments;
[0041] Perform synchronous recognition on multiple sub-consultation segments.
[0042] Optionally, divide the overall consultation image into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture image and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person, and further include:
[0043] Define corresponding posture images and voice segments based on each sub-consultation segment;
[0044] Associate multiple sub-consultation segments;
[0045] Output the Chinese information of each sub-consultation segment based on the synthesis of the posture image and the voice segment. At this time, the posture image and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person.
[0046] Optionally, match the corresponding recognition and learning model according to the overall consultation image, the Chinese information, and the action time nodes of each person, define an abnormal node according to the recognition of the recognition and learning model and the overall consultation image, and perform local optimization on the abnormal node to improve the overall consultation image, including:
[0047] Collect the overall consultation image and the corresponding Chinese information;
[0048] Freeze the overall consultation picture, Chinese information, and the action time nodes of each person;
[0049] Associate the overall consultation picture, Chinese information, and the action time nodes of each person.
[0050] Optionally, match the corresponding recognition and learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, define abnormal nodes based on the recognition and learning model and the recognition of the overall consultation picture, and perform local optimization for the abnormal nodes to improve the overall consultation picture. It further includes:
[0051] Match the corresponding recognition and learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person;
[0052] Define abnormal nodes based on the recognition and learning model and the recognition of the overall consultation picture;
[0053] Trigger picture tracing according to the abnormal nodes to perform local optimization for the abnormal nodes to improve the overall consultation picture.
[0054] In addition, the embodiment of the present invention further provides a voice translation system for the overall consultation picture. The voice translation system for the overall consultation picture includes:
[0055] The primary consultation module is used for the rescue personnel to conduct a primary consultation on the wounded in the first consultation space of the ambulance and collect the preliminary results of the primary consultation;
[0056] The secondary consultation module is used for remotely interacting with the first consultation space of the ambulance and the second consultation space of the hospital, and defining the secondary consultation results according to the medical staff in the second consultation space and the preliminary results of the primary consultation;
[0057] The interaction module is used for defining the interaction time period based on the time point where the preliminary results of the primary consultation are located and the time point where the secondary consultation results are located;
[0058] The picture module is used for constructing the overall consultation picture according to the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space;
[0059] The translation module is used for dividing the overall consultation picture into multiple sub-consultation segments, outputting the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture picture and the voice segment are dynamically matched, and the Chinese information described by each person is translated based on the association of the actions, lip shapes, and corresponding voices of each person;
[0060] An optimization module, configured to match a corresponding recognition and learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, define an abnormal node based on the recognition of the recognition and learning model and the overall consultation picture, and perform local optimization on the abnormal node to improve the overall consultation picture.
[0061] In an embodiment of the present invention, by the method in the embodiment of the present invention, the ambulance personnel perform a primary consultation on the wounded in the first consultation space of the ambulance and collect the preliminary results of the primary consultation; remotely interact with the first consultation space of the ambulance and the second consultation space of the hospital, and define the secondary consultation results according to the medical staff in the second consultation space and the preliminary results of the primary consultation; define an interaction time period based on the time point where the preliminary results of the primary consultation are located and the time point where the secondary consultation results are located; construct an overall consultation picture according to the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, which comprehensively considers the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, and multi-dimensionally controls the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, ensuring the accuracy of the overall consultation picture.
[0062] Further, divide the overall consultation picture into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture picture and the voice segment are dynamically matched, and the Chinese information described by each person is translated based on the association of the actions, lip shapes, and corresponding voices of each person; match the corresponding recognition and learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, define an abnormal node based on the recognition of the recognition and learning model and the overall consultation picture, and perform local optimization on the abnormal node to improve the overall consultation picture. Multiple mechanisms are performed for the output of the overall consultation picture and the recognition of the overall consultation picture, so as to achieve multi-level control of the overall consultation picture, ensure the high matching degree of the Chinese information and each segment in the overall consultation picture, and improve the voice translation effect of the overall consultation picture. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0064] Figure 1 It is a schematic flowchart of the voice translation method for the overall consultation picture in the embodiment of the present invention;
[0065] Figure 2 It is a schematic flowchart of S11 in the voice translation method for the overall consultation picture in the embodiments of the present invention;
[0066] Figure 3 It is a schematic flowchart of S12 in the voice translation method for the overall consultation picture in the embodiments of the present invention;
[0067] Figure 4 It is a schematic flowchart of S13 in the voice translation method for the overall consultation picture in the embodiments of the present invention;
[0068] Figure 5 It is a schematic flowchart of S14 in the voice translation method for the overall consultation picture in the embodiments of the present invention;
[0069] Figure 6 It is a schematic flowchart of S15 in the voice translation method for the overall consultation picture in the embodiments of the present invention;
[0070] Figure 7 It is a schematic flowchart of S16 in the voice translation method for the overall consultation picture in the embodiments of the present invention;
[0071] Figure 8 It is a schematic diagram of the structural composition of the voice translation system for the overall consultation picture in the embodiments of the present invention;
[0072] Figure 9 It is a hardware diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0073] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0074] Please refer to Figures 1 to 9 , a voice translation method for the overall consultation picture, which is applied to the voice translation scenario of the overall consultation picture; the voice translation method for the overall consultation picture includes:
[0075] Step S11: The rescue personnel conduct a primary consultation on the wounded based on the first consultation space of the ambulance and collect the preliminary results of the primary consultation;
[0076] Step S12: Remotely interact with the first consultation space of the ambulance and the second consultation space of the hospital, and define the secondary consultation results according to the medical staff in the second consultation space and the preliminary results of the primary consultation;
[0077] Step S13: Define an interaction time period based on the time point of the preliminary result of the first-level consultation and the time point of the second-level consultation result;
[0078] Step S14: Construct an overall consultation picture according to the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space;
[0079] Step S15: Divide the overall consultation picture into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture picture and the voice segment are dynamically matched, and the Chinese information described by each person is translated based on the association between the actions, lip shapes, and corresponding voices of each person;
[0080] Step S16: Match the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person. Define an abnormal node according to the recognition of the recognition learning model and the overall consultation picture, and perform local optimization on the abnormal node to improve the overall consultation picture.
[0081] In the embodiment of the present invention, through the method in the embodiment of the present invention, the rescue personnel conduct a first-level consultation on the wounded based on the first consultation space of the ambulance and collect the preliminary result of the first-level consultation; remotely interact with the first consultation space of the ambulance and the second consultation space of the hospital, and define the second-level consultation result according to the medical staff in the second consultation space and the preliminary result of the first-level consultation; define an interaction time period based on the time point of the preliminary result of the first-level consultation and the time point of the second-level consultation result; construct an overall consultation picture according to the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space, which takes into account the overall consideration of the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space, and multi-dimensionally controls the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space, ensuring the accuracy of the overall consultation picture.
[0082] Furthermore, divide the overall consultation picture into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture picture and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person; match the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, define abnormal nodes according to the recognition learning model and the recognition of the overall consultation picture, and perform local optimization for the abnormal nodes to improve the overall consultation picture. For the output of the overall consultation picture and the recognition of the overall consultation picture, multiple mechanisms are carried out to facilitate the multi-level control of the overall consultation picture, ensure the high matching of the Chinese information and each segment in the overall consultation picture, and improve the voice translation effect of the overall consultation picture.
[0083] Reference Figure 2 , in step S11, the ambulance personnel conduct a primary consultation on the wounded based on the first consultation space of the ambulance, and collect the preliminary results of the primary consultation;
[0084] In the specific implementation process of the present invention, the specific steps may be:
[0085] S111: Collect the arrival signal of the ambulance, locate the ambulance position according to the analysis of the arrival signal of the ambulance, and collect the accident segment according to the location detection of the ambulance position;
[0086] S112: Define the accident type based on the accident segment and the accident learning model, and define the preliminary injury of the wounded according to the accident type, the injured segment of the wounded, and the duration, and transmit the preliminary injury of the wounded to the first consultation space of the ambulance;
[0087] S113: Define the injury form according to the preliminary injury of the wounded and the actual injured image of the wounded, and perform preliminary treatment according to the injury form and the rescue tools available in the first consultation space. The ambulance personnel conduct a primary consultation on the wounded based on the first consultation space of the ambulance;
[0088] S114: Collect the preliminary results of the primary consultation.
[0089] In the embodiment of the present application, by collecting the arrival signal of the ambulance, locating the ambulance position according to the analysis of the arrival signal of the ambulance, and collecting the accident segment according to the location detection of the ambulance position, the accident segment is introduced, and the location detection of the ambulance position is realized, so as to further know the accident segment and realize the subsequent processing of the accident segment.
[0090] At this time, an accident type is defined based on the accident segment and the accident learning model, and a preliminary injury condition of the wounded is defined according to the accident type, the injury segment of the wounded, and the duration, taking into account the accident type, the injury segment of the wounded, and the duration as a whole, achieving multi-dimensional control of the accident type, the injury segment of the wounded, and the duration, so as to ensure the accuracy of the preliminary injury condition of the wounded. At the same time, the preliminary injury condition of the wounded is transmitted to the first consultation space of the ambulance.
[0091] Therefore, an injury form is defined according to the preliminary injury condition of the wounded and the actual injury image of the wounded, and preliminary treatment is carried out according to the injury form and the rescue tools available in the first consultation space. The rescue personnel conduct a first-level consultation on the wounded based on the first consultation space of the ambulance, so as to conduct a first-level consultation on the first consultation space of the ambulance, making full use of the first consultation space and reflecting the timely response to the preliminary injury condition of the wounded, and then collecting the preliminary results of the first-level consultation.
[0092] Reference Figure 3 , in step S12, the first consultation space of the remote interactive ambulance and the second consultation space of the hospital are interacted, and a second-level consultation result is defined according to the medical staff in the second consultation space and the preliminary results of the first-level consultation;
[0093] In the specific implementation process of the present invention, the specific steps may be:
[0094] S121: Freeze the first consultation space of the ambulance;
[0095] S122: Collect the second consultation space of the hospital;
[0096] S123: Conduct remote interaction for the second consultation space of the hospital and the first consultation space of the ambulance;
[0097] S124: Freeze the preliminary results of the first-level consultation; during the remote interaction process, the medical staff in the second consultation space make a preliminary judgment on the preliminary results of the first-level consultation and output the corresponding judgment results;
[0098] S125: Associate the judgment result, the preliminary results of the first-level consultation, and the injury form of the wounded; conduct multiple interactions on the judgment result, the preliminary results of the first-level consultation, and the injury form of the wounded, and define the corresponding second-level consultation result based on the judgment result, the preliminary results of the first-level consultation, and the injury form of the wounded.
[0099] In the embodiments of the present application, the first consultation space of the stationary ambulance is captured, and at the same time, the second consultation space of the hospital is captured. Remote interaction is performed on the second consultation space of the hospital and the first consultation space of the ambulance, realizing the remote interaction between the second consultation space and the second consultation space, reflecting the real-time association between the hospital and the ambulance, and ensuring the timely response of the second consultation space and the second consultation space.
[0100] At this time, the preliminary results of the first-level consultation are frozen; during the remote interaction process, the medical staff in the second consultation space make a preliminary judgment on the preliminary results of the first-level consultation and output the corresponding judgment results, so as to further control the judgment results.
[0101] Therefore, the judgment result, the preliminary result of the first-level consultation, and the injury form of the wounded are associated; multiple interactions are performed on the judgment result, the preliminary result of the first-level consultation, and the injury form of the wounded, and the corresponding second-level consultation result is defined based on the judgment result, the preliminary result of the first-level consultation, and the injury form of the wounded, accommodating the overall consideration of the judgment result, the preliminary result of the first-level consultation, and the injury form of the wounded, realizing the multi-dimensional control of the judgment result, the preliminary result of the first-level consultation, and the injury form of the wounded, and ensuring the rationality and accuracy of the second-level consultation result.
[0102] Reference Figure 4 In step S13, an interaction time period is defined based on the time point where the preliminary result of the first-level consultation is located and the time point where the second-level consultation result is located;
[0103] In the specific implementation process of the present invention, the specific steps may be:
[0104] S131: Freeze the preliminary result of the first-level consultation and the second-level consultation result;
[0105] S132: Collect the time point where the preliminary result of the first-level consultation is located, and at the same time, collect the time point where the second-level consultation result is located;
[0106] S133: Associate the time point where the preliminary result of the first-level consultation is located and the time point where the second-level consultation result is located;
[0107] S134: Define an interaction time period according to the time point where the preliminary result of the first-level consultation is located and the time point where the second-level consultation result is located;
[0108] S135: Define an abnormal time based on the traversal of the interaction time period, define the corresponding abnormal picture based on the traceability of the abnormal time, and optimize the abnormal time according to the intelligent recognition of the abnormal picture to improve the interaction time period.
[0109] In the embodiments of the present application, freeze the preliminary results of the first-level consultation and the results of the second-level consultation; collect the time point where the preliminary results of the first-level consultation are located, and at the same time, collect the time point where the results of the second-level consultation are located. Introduce the time point where the preliminary results of the first-level consultation are located and the time point where the results of the second-level consultation are located, so as to associate the time point where the preliminary results of the first-level consultation are located and the time point where the results of the second-level consultation are located, and realize the further processing of the time point where the preliminary results of the first-level consultation are located and the time point where the results of the second-level consultation are located.
[0110] Therefore, define the interaction time period according to the time point where the preliminary results of the first-level consultation are located and the time point where the results of the second-level consultation are located. Introduce the interaction time period, realize the precise control of the interaction time period. At the same time, define the abnormal time based on the traversal of the interaction time period, define the corresponding abnormal screen based on the traceability of the abnormal time, and optimize the abnormal time according to the intelligent recognition of the abnormal screen to improve the interaction time period, so as to facilitate the traceability and optimization of the abnormal screen, thereby controlling the abnormal time and realizing the further improvement of the interaction time period.
[0111] Reference Figure 5 , S14: Construct the overall consultation screen according to the interaction time period, the screen captured by the first consultation space, and the screen captured by the second consultation space;
[0112] In the specific implementation process of the present invention, the specific steps can be:
[0113] S141: Freeze the interaction time period;
[0114] S142: Associate the interaction time period, the screen captured by the first consultation space, and the screen captured by the second consultation space;
[0115] S143: Define the first sub-screen based on the interaction time period and the screen captured by the first consultation space;
[0116] S144: Define the second sub-screen based on the interaction time period and the screen captured by the second consultation space, and define the transition screen based on the screen matching of the second sub-screen and the first sub-screen;
[0117] S145: Construct the overall consultation screen based on the first sub-screen, the second sub-screen, and the transition screen.
[0118] In an embodiment of the present application, the rescue personnel conduct a first-level consultation on the wounded based on the first consultation space of the ambulance and collect the preliminary results of the first-level consultation; remotely interact with the first consultation space of the ambulance and the second consultation space of the hospital, and define the second-level consultation results according to the medical staff in the second consultation space and the preliminary results of the first-level consultation; define the interaction time period based on the time point where the preliminary results of the first-level consultation are located and the time point where the second-level consultation results are located; construct the overall consultation picture according to the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, which takes into account the overall consideration of the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, and multi-dimensionally controls the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, ensuring the accuracy of the overall consultation picture.
[0119] At this time, freeze the interaction time period; associate the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, realizing multiple interactions of the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, and ensuring multi-dimensional control of the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space.
[0120] Therefore, define the first sub-picture based on the interaction time period and the pictures taken by the first consultation space; define the second sub-picture based on the interaction time period and the pictures taken by the second consultation space, and define the transition picture based on the picture matching of the second sub-picture and the first sub-picture; construct the overall consultation picture based on the first sub-picture, the second sub-picture, and the transition picture, which takes into account the overall consideration of the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, and multi-dimensionally controls the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, ensuring the accuracy of the overall consultation picture.
[0121] Reference Figure 6 , S15: Divide the overall consultation picture into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture picture and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person;
[0122] In the specific implementation process of the present invention, the specific steps may be:
[0123] S151: Freeze the overall consultation picture;
[0124] S152: Divide the overall consultation picture into multiple sub-consultation segments;
[0125] S153: Synchronously recognize multiple sub-consultation segments;
[0126] S154: Define corresponding gesture screens and voice segments based on each sub-consultation segment;
[0127] S155: Associate multiple sub-consultation segments;
[0128] S156: Output the Chinese information of the sub-consultation segment based on the synthesis of the gesture screen and the voice segment. At this time, the gesture screen and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person.
[0129] In the embodiment of the present application, freeze the overall consultation screen and further process the overall consultation screen. At this time, divide the overall consultation screen into multiple sub-consultation segments, introduce multiple sub-consultation segments, and thus synchronously recognize multiple sub-consultation segments.
[0130] Therefore, define corresponding gesture screens and voice segments based on each sub-consultation segment; associate multiple sub-consultation segments; output the Chinese information of the sub-consultation segment based on the synthesis of the gesture screen and the voice segment. At this time, the gesture screen and the voice segment are dynamically matched, achieving accurate matching of the gesture screen and the voice segment, ensuring the consistency between the gesture screen and the voice segment. At the same time, it is translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person, achieving multi-dimensional control of the actions, lip shapes, and corresponding voices of each person, and ensuring the accuracy and correspondence of the Chinese information described by each person.
[0131] Reference Figure 7 , S16: Match the corresponding recognition and learning model according to the overall consultation screen, Chinese information, and the action time nodes of each person. Define abnormal nodes based on the recognition of the recognition and learning model and the overall consultation screen, and perform local optimization for the abnormal nodes to improve the overall consultation screen;
[0132] In the specific implementation process of the present invention, the specific steps can be:
[0133] S161: Collect the overall consultation screen and the corresponding Chinese information;
[0134] S162: Freeze the overall consultation screen, Chinese information, and the action time nodes of each person;
[0135] S163: Associate the overall consultation screen, Chinese information, and the action time nodes of each person;
[0136] S164: Match the corresponding recognition and learning model according to the overall consultation screen, Chinese information, and the action time nodes of each person;
[0137] S165: Define abnormal nodes based on the recognition learning model and the recognition of the overall consultation picture;
[0138] S166: Trigger picture tracing according to the abnormal nodes to perform local optimization for the abnormal nodes to improve the overall consultation picture.
[0139] In the specific implementation process of the present invention, multiple sub-consultation segments are divided based on the overall consultation picture, and the Chinese information of each sub-consultation segment is output based on each sub-consultation segment. At this time, the posture picture and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes and corresponding voices of each person; according to the overall consultation picture, Chinese information, and the action time nodes of each person, the corresponding recognition learning model is matched, and abnormal nodes are defined according to the recognition learning model and the recognition of the overall consultation picture, and local optimization is performed for the abnormal nodes to improve the overall consultation picture. Multiple mechanisms are performed for the output of the overall consultation picture and the recognition of the overall consultation picture, so as to realize multi-level control of the overall consultation picture, ensure the high matching degree of the Chinese information and each segment in the overall consultation picture, and improve the voice translation effect of the overall consultation picture.
[0140] At this time, the overall consultation picture and the corresponding Chinese information are collected, so as to freeze the overall consultation picture, Chinese information, and the action time nodes of each person, and then associate the overall consultation picture, Chinese information, and the action time nodes of each person.
[0141] Therefore, match the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person; define abnormal nodes according to the recognition learning model and the recognition of the overall consultation picture; trigger picture tracing according to the abnormal nodes to perform local optimization for the abnormal nodes to improve the overall consultation picture. Optionally, the recognition learning model is trained based on past training data.
[0142] Furthermore, multiple mechanisms are performed for the output of the overall consultation picture and the recognition of the overall consultation picture, so as to realize multi-level control of the overall consultation picture, ensure the high matching degree of the Chinese information and each segment in the overall consultation picture, and improve the voice translation effect of the overall consultation picture.
[0143] In an embodiment of the present invention, through the method in the embodiment of the present invention, rescue personnel conduct a first-level consultation on the wounded based on the first consultation space of the ambulance and collect the preliminary results of the first-level consultation; remotely interact with the first consultation space of the ambulance and the second consultation space of the hospital, and define the second-level consultation results according to the medical staff in the second consultation space and the preliminary results of the first-level consultation; define the interaction time period based on the time point where the preliminary results of the first-level consultation are located and the time point where the second-level consultation results are located; construct the overall consultation picture according to the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, taking into account the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space in a compatible manner, and multi-dimensionally control the interaction time period, the pictures taken by the first consultation space, and the pictures taken by the second consultation space, ensuring the accuracy of the overall consultation picture.
[0144] Further, divide the overall consultation picture into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the gesture picture and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip movements, and corresponding voices of each person; match the corresponding recognition learning model according to the overall consultation picture, the Chinese information, and the action time nodes of each person, define the abnormal nodes according to the recognition of the recognition learning model and the overall consultation picture, and perform local optimization on the abnormal nodes to improve the overall consultation picture. For the output of the overall consultation picture and the recognition of the overall consultation picture, multiple mechanisms are used to facilitate the multi-level control of the overall consultation picture, ensure the high matching degree between the Chinese information and each segment of the overall consultation picture, and improve the voice translation effect of the overall consultation picture.
[0145] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the voice translation system for the overall consultation picture in the embodiment of the present invention.
[0146] As Figure 8 shown, a voice translation system for an overall consultation picture, the voice translation system for the overall consultation picture includes:
[0147] A first-level consultation module 21, configured to enable rescue personnel to conduct a first-level consultation on the wounded based on the first consultation space of the ambulance and collect the preliminary results of the first-level consultation;
[0148] A second-level consultation module 22, configured to remotely interact with the first consultation space of the ambulance and the second consultation space of the hospital, and define the second-level consultation results according to the medical staff in the second consultation space and the preliminary results of the first-level consultation;
[0149] An interaction module 23, configured to define an interaction time period based on the time point where the preliminary result of the first-level consultation is located and the time point where the second-level consultation result is located;
[0150] A screen module 24, configured to construct an overall consultation screen according to the interaction time period, the screen captured by the first consultation space, and the screen captured by the second consultation space;
[0151] A translation module 25, configured to divide a plurality of sub-consultation segments based on the overall consultation screen, output Chinese information of the sub-consultation segment based on each sub-consultation segment. At this time, the posture screen and the voice segment are dynamically matched, and are translated into the Chinese information described by each person based on the association of the actions, lip shapes, and corresponding voices of each person;
[0152] An optimization module 26, configured to match a corresponding recognition learning model according to the overall consultation screen, the Chinese information, and the action time nodes of each person, define an abnormal node according to the recognition learning model and the recognition of the overall consultation screen, and perform local optimization on the abnormal node to improve the overall consultation screen.
[0153] Please refer to Figure 9 , and hereinafter refer to Figure 9 to describe the electronic intelligent device 40 according to this embodiment of the present invention. Figure 9 The shown electronic intelligent device 40 is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0154] As Figure 9 shown, the electronic intelligent device 40 is presented in the form of a general computing intelligent device. The components of the electronic intelligent device 40 may include, but are not limited to: at least one of the above-mentioned processing units 41, at least one of the above-mentioned storage units 42, and a bus 43 connecting different system components (including the storage unit 42 and the processing unit 41).
[0155] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 41, so that the processing unit 41 executes the steps according to various exemplary embodiments of the present invention described in the "Embodiment Method" part of this specification.
[0156] The storage unit 42 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 421 and / or a cache storage unit 422, and may further include a read-only storage unit (ROM) 423.
[0157] The storage unit 42 may also include a program / utilities 424 having a set (at least one) of program modules 425. Such program modules 425 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0158] The bus 43 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0159] The electronic intelligent device 40 may also communicate with one or more external intelligent devices (such as a keyboard, a pointing intelligent device, a Bluetooth intelligent device, etc.), may also communicate with one or more intelligent devices that enable a user to interact with the electronic intelligent device 40, and / or may communicate with any intelligent device (such as a router, a modem, etc.) that enables the electronic intelligent device 40 to communicate with one or more other computing intelligent devices. Such communication may be through the input / output (I / O) interface 44. Also, the electronic intelligent device 40 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 45. As Figure 9 shown, the network adapter 45 communicates with other modules of the electronic intelligent device 40 through the bus 43. It should be understood that although Figure 9 not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic intelligent device 40, including but not limited to: microcode, intelligent device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup planning systems, etc.
[0160] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing intelligent device (which may be a personal computer, a server, a terminal device, or a network intelligent device, etc.) to execute the method according to the embodiments of the present disclosure.
[0161] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc. Moreover, it stores computer program instructions, and when the computer program instructions are executed by a computer, the computer is made to execute the method according to the above.
[0162] In addition, the above has introduced in detail the voice translation method and system for the overall consultation picture provided by the embodiments of the present invention. Specific examples should have been used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for voice translation of the overall consultation picture, characterized in that: Applied to the speech translation scenario of the overall picture of the consultation; The method for voice translation of the overall consultation picture includes: The rescuer conducts a primary consultation on the injured person based on the primary consultation space of the ambulance and collects the preliminary results of the primary consultation; The first consultation space of the remote interactive ambulance and the second consultation space of the hospital define the secondary consultation results according to the medical staff in the second consultation space and the preliminary results of the first consultation; The interaction time period is defined based on the time point of the primary consultation's preliminary results and the time point of the secondary consultation's results; Constructing an overall consultation picture according to the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space; Divide the overall consultation image into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on the sub-consultation segment. At this time, the posture image and the voice sound segment are dynamically matched, and translated into the Chinese information described by each person based on the association between the action, mouth shape and corresponding sound of each person; According to the overall picture of the consultation, Chinese information, and the action time nodes of each person, the corresponding recognition learning model is matched. According to the recognition learning model and the recognition of the overall picture of the consultation, the abnormal node is defined, and local optimization is performed on the abnormal node to improve the overall picture of the consultation.
2. The method for voice translation of the overall consultation picture according to claim 1, characterized in that: The rescuer conducts a primary consultation on the injured person based on the primary consultation space of the ambulance, and collects preliminary results of the primary consultation, including: Collect the arrival signal of the ambulance, locate the rescue position according to the analysis of the arrival signal of the ambulance, and collect the accident clip according to the positioning detection of the rescue position; The accident type is defined based on the accident fragment and the accident learning model, and the initial injury of the injured person is defined according to the accident type, the injured fragment and the duration, and the initial injury of the injured person is transmitted to the first consultation space of the ambulance; The injury morphology is defined according to the initial injury of the injured and the actual injury image of the injured, and preliminary treatment is performed according to the injury morphology and the rescue tools available in the first consultation space. The rescue personnel conduct a primary consultation on the injured based on the first consultation space of the ambulance; Collect preliminary results of the first-level consultation.
3. The method for voice translation of the overall consultation picture according to claim 2, characterized in that: The first consultation space of the remote interactive ambulance and the second consultation space of the hospital define the secondary consultation results according to the medical staff in the second consultation space and the preliminary results of the primary consultation, including: Freeze the first consultation space in the ambulance; The second consultation space of the collection hospital; Remote interaction with the second consultation space in the hospital and the first consultation space in the ambulance; Freeze the preliminary results of the first-level consultation; during the remote interaction, the medical staff in the second consultation space make a preliminary judgment on the preliminary results of the first-level consultation and output the corresponding judgment results; The judgment result, the preliminary result of the first-level consultation and the injury form of the injured person are associated; multiple interactions are performed on the judgment result, the preliminary result of the first-level consultation and the injury form of the injured person, and a corresponding second-level consultation result is defined based on the judgment result, the preliminary result of the first-level consultation and the injury form of the injured person.
4. The method for voice translation of the overall consultation picture according to claim 3, characterized in that: The interaction time period is defined based on the time point of the preliminary result of the first-level consultation and the time point of the secondary consultation result, including: Finalize the preliminary results of the first-level consultation and the results of the second-level consultation; The time point at which the preliminary results of the first-level consultation are collected, and at the same time, the time point at which the results of the second-level consultation are collected; The time point of the preliminary results of the first-level consultation and the time point of the results of the second-level consultation are related; The interaction time period is defined according to the time point of the primary consultation's preliminary results and the time point of the secondary consultation's results; The abnormal time is defined based on the traversal of the interactive time period, the corresponding abnormal screen is defined based on the tracing of the abnormal time, and the abnormal time is optimized according to the intelligent identification of the abnormal screen to improve the interactive time period.
5. The method for voice translation of the overall consultation picture according to claim 4, characterized in that: The step of constructing an overall consultation picture according to the interactive time period, the picture taken by the first consultation space, and the picture taken by the second consultation space includes: Freeze the interaction time period; Associating the interaction time period, the images captured in the first consultation space, and the images captured in the second consultation space; defining a first sub-image based on the interactive time period and the image captured in the first consultation space; A second sub-picture is defined based on the interactive time period and the pictures taken in the second consultation space, and a transition picture is defined based on the picture matching between the second sub-picture and the first sub-picture; The overall consultation picture is constructed based on the first sub-picture, the second sub-picture and the transition picture.
6. The method for voice translation of the overall consultation picture according to claim 5, characterized in that: The method of dividing the consultation into multiple sub-consultation segments based on the overall consultation screen and outputting the Chinese information of each sub-consultation segment based on each sub-consultation segment, wherein the posture screen and the voice sound segment are dynamically matched and translated into the Chinese information described by each person based on the association between the action, mouth shape and corresponding sound of each person, includes: Freeze the overall picture of the consultation; Divide the consultation into multiple sub-consultation segments based on the overall picture of the consultation; Simultaneous recognition of multiple sub-consultation segments.
7. The method for voice translation of the overall consultation picture according to claim 6, characterized in that: The method further includes dividing the consultation image into multiple sub-consultation segments based on the overall consultation image, and outputting Chinese information of each sub-consultation segment based on each sub-consultation segment. At this time, the posture image and the voice sound segment are dynamically matched, and translated into Chinese information described by each person based on the association between the action, mouth shape and corresponding sound of each person. Based on each sub-consultation segment, corresponding posture images and voice sound segments are defined; Link multiple sub-consultation fragments; The Chinese information of the sub-consultation segment is output based on the synthesis of the posture image and the voice sound segment. At this time, the posture image and the voice sound segment are dynamically matched and translated into the Chinese information described by each person based on the association between each person's actions, mouth shape and corresponding sound.
8. The method for voice translation of the overall consultation picture according to claim 7, characterized in that: The method matches the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, defines an abnormal node according to the recognition learning model and the recognition of the overall consultation picture, and performs local optimization on the abnormal node to improve the overall consultation picture, including: Collect the overall picture of the consultation and the corresponding Chinese information; Freeze the overall consultation picture, Chinese information, and the action time nodes of each person; Associate the overall consultation picture, Chinese information, and the action time nodes of each person.
9. The method for voice translation of the overall consultation picture according to claim 8, characterized in that: The method further includes matching the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, defining abnormal nodes according to the recognition learning model and the recognition of the overall consultation picture, and performing local optimization on the abnormal nodes to improve the overall consultation picture. Match the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person; The abnormal nodes are defined based on the recognition learning model and the recognition of the overall picture of the consultation; Trigger screen tracing based on abnormal nodes to perform local optimization on the abnormal nodes to improve the overall picture of the consultation.
10. A voice translation system for the overall consultation picture, characterized in that: The speech translation system of the overall consultation picture is applied to the speech translation method of the overall consultation picture as claimed in any one of claims 1 to 9, and the speech translation system of the overall consultation picture includes: The first-level consultation module is used by the rescue personnel to conduct a first-level consultation on the injured based on the first consultation space of the ambulance and collect the preliminary results of the first-level consultation; The secondary consultation module is used for remotely interacting with the first consultation space of the ambulance and the second consultation space of the hospital, and defines the secondary consultation results according to the medical staff in the second consultation space and the preliminary results of the first consultation; An interaction module, used to define an interaction time period based on a time point of a preliminary result of the primary consultation and a time point of a result of the secondary consultation; A picture module, used to construct an overall consultation picture according to the interaction time period, the pictures taken in the first consultation space, and the pictures taken in the second consultation space; The translation module is used to divide the overall consultation image into multiple sub-consultation segments, and output the Chinese information of each sub-consultation segment based on the sub-consultation segment. At this time, the posture image and the voice sound segment are dynamically matched, and translated into the Chinese information described by each person based on the association of each person's action, mouth shape and corresponding sound; The optimization module is used to match the corresponding recognition learning model according to the overall consultation picture, Chinese information, and the action time nodes of each person, define abnormal nodes based on the recognition learning model and the recognition of the overall consultation picture, and perform local optimization on the abnormal nodes to improve the overall consultation picture.
Citation Information
Patent Citations
Audio and video playing method and device, electronic equipment and storage medium
CN112437336A
5G interconnection transfer system and method applied to medical first aid
CN117637190A
Structured data synchronization method and system
CN118170852A
Intelligent operation teaching and monitoring early warning method and system
CN118865204A
Remote conference real-time voice recognition optimization method and system based on visual context
CN119207406A