Acoustic wave annotation training method, system and device for physical change phenomena

Through the sonic wave labeling training method, the physical changes are automatically marked by sound wave signals, which solves the problem of inaccurate training under the influence of light and angle in the existing technology, and achieves high-precision analysis results and automated training.

WO2025180438A1PCT designated stage Publication Date: 2025-09-04CHEN CHIAHUNG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/079497
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2025-02-27
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

The prior art is susceptible to factors such as light and angle when analyzing physical changes, resulting in inaccurate training results, especially in the subtle phenomena of physical changes, which are difficult to achieve high-precision image training.

Method used

The acoustic wave labeling training method is adopted to obtain physical changes through the acquisition unit, the transmitting end generates and transmits acoustic wave signals, the receiving end performs operation analysis, and the operation processing unit performs multimodal operation analysis based on the changes in the acoustic wave signals, producing high-precision analysis results.

Benefits of technology

Automatic labeling training is realized, the accuracy of analysis results is improved, manual intervention is reduced, and the accuracy of multimodal operation analysis is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025079497_04092025_PF_FP_ABST
    Figure CN2025079497_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention are an acoustic wave annotation training method, system and device for physical change phenomena. The acoustic wave annotation training method includes the steps of: an acquisition unit acquiring a physical change phenomenon; on the basis of the physical change phenomenon, a transmitting end generating a corresponding acoustic wave signal and transmitting same; a receiving end receiving the acoustic wave signal and performing operational analysis on the basis of a change in the acoustic wave signal, so as to generate corresponding physical change information; and an operational processing unit receiving the physical change information and annotating the physical change phenomenon on the basis of the physical change information, so as to execute multimodal operational analysis to generate an analysis result. The method is executed by means of the system / device to implement automatic annotation, thereby establishing high-precision multimodal analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Sound wave annotation training method, system and equipment applied to physical change phenomena Technical Field

[0001] The present invention belongs to the field of labeling training methods, systems and equipment thereof, and in particular relates to a sound wave labeling training method, system and equipment thereof applied to physical change phenomena. Background Art

[0002] Although image training is relatively mature in existing technologies, it still has its limitations. It is particularly susceptible to data bias, especially in physical changes. Different lighting and angles may cause visual errors. For example, boiling water at 100°C may appear similar to ice in an image under different lighting and angles, resulting in inaccurate training results. In addition, changes in physical phenomena such as airflow speed, humidity, light, and fluctuations may result in very subtle differences in images, greatly increasing the difficulty of image training.

[0003] To this end, the present invention provides a sound wave annotation training method, system, and equipment for physical change phenomena, which automatically convert and perform annotation based on changes in sound wave signals. This not only significantly improves the accuracy of analysis results, but also effectively solves the problem that changes in physical phenomena in traditional images are difficult to learn and train. Summary of the Invention

[0004] The main purpose of the present invention is to provide an acoustic wave annotation training method for physical change phenomena, which transmits physical change phenomena through acoustic wave signals, thereby automatically annotating and training physical change phenomena to establish highly accurate analysis results.

[0005] Another object of the present invention is to provide an acoustic wave annotation training system for physical change phenomena, which uses a conversion unit to generate and emit acoustic wave signals based on the physical change phenomenon, thereby transmitting its physical change information, so that the computing processing unit can be trained to annotate the corresponding physical change phenomenon through this physical change information to achieve automatic annotation.

[0006] Another object of the present invention is to provide an acoustic wave annotation training device for physical change phenomena. The detection device transmits the physical change phenomenon through acoustic wave signals, allowing the computing device to perform subsequent annotation and training to obtain highly accurate analysis results.

[0007] In order to achieve the above-mentioned purpose, one embodiment of the present invention discloses a sound wave annotation training method applied to physical change phenomena, the steps including: using a capture unit to acquire a physical change phenomenon; a transmitting end to generate and transmit a corresponding sound wave signal based on the physical change phenomenon; using a receiving end to receive and perform computational analysis based on the change of the sound wave signal to generate corresponding physical change information; and a computational processing unit to receive and annotate the physical change phenomenon based on the physical change information to perform multimodal computational analysis and generate an analysis result.

[0008] In a preferred embodiment, in a step in which a processing unit labels the physical change phenomenon according to the physical change information to perform multimodal computational analysis and generate an analysis result, the processing unit receives and performs computational processing based on the physical change information to generate corresponding semantic information, so that the processing unit labels the physical change phenomenon according to the semantic information.

[0009] In a preferred embodiment, the type of the acquired physical change phenomenon is selected from detected electronic signals, images, audio and video, voice, text, or a combination of any two or more of the above.

[0010] In a preferred embodiment, the physical change phenomenon is selected from spatial position change, temperature change, elastic change, electromagnetic change, light change, wave change, vibration change, mass change, gas flow rate change, humidity change or a combination of any two or more of the above.

[0011] In order to achieve the other purpose mentioned above, one embodiment of the present invention discloses a sound wave annotation training system applied to physical change phenomena, comprising: an acquisition unit for acquiring a physical change phenomenon; a conversion unit comprising a transmitting end and a receiving end, the transmitting end being used to generate and transmit a corresponding sound wave signal according to the physical change phenomenon, and the receiving end being used to receive and perform computational analysis based on the change of the sound wave signal to generate corresponding physical change information; and an computational processing unit being signal-connected to the acquisition unit and the conversion unit respectively, for receiving and annotating the physical change phenomenon according to the physical change information and performing multimodal computational analysis to generate an analysis result.

[0012] In a preferred embodiment, the capture unit is selected from an image capture device, a voice capture device, an input device, a sensing device, or a combination of any two or more of the above.

[0013] In order to achieve the above-mentioned further purpose, one embodiment of the present invention discloses a sound wave annotation training device applied to physical change phenomena, comprising: a detection device for generating and emitting a sound wave signal according to a physical change phenomenon; and a computing device, signal-connected to the detection device, for receiving and performing computational analysis based on the change of the sound wave signal, generating corresponding physical change information, and annotating the physical change phenomenon according to the physical change information to perform multimodal computational analysis and generate an analysis result.

[0014] In a preferred embodiment, the detection device includes a sensing device for detecting the physical change phenomenon and generating the physical change information.

[0015] In order to achieve another of the above-mentioned purposes, one embodiment of the present invention discloses an intelligent automation system, comprising: an automation device; a detection device, signal-connected to the automation device, for acquiring image information, and generating and emitting a corresponding acoustic wave signal according to a physical change phenomenon in the image information; a computing device, signal-connected to the detection device, for receiving and performing computational analysis based on the change of the acoustic wave signal to generate corresponding physical change information, and marking the image information according to the physical change information to perform multimodal computational analysis and generate an analysis result; and a decision-making device, signal-connected to the automation device and the computing device, respectively, for predicting and outputting decision information based on the analysis result, so that the automation device operates according to the decision information.

[0016] In a preferred embodiment, the automated equipment is selected from a robotic arm, a self-propelled device, a drone or a robot.

[0017] The beneficial effect of the present invention lies in automated annotation training, thereby establishing highly accurate analysis results and improving the accuracy of prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a flow chart of a method according to an embodiment of the present invention;

[0019] FIG2A is a system schematic diagram of an embodiment of the present invention;

[0020] FIG2B is a schematic diagram of a system operation according to an embodiment of the present invention;

[0021] FIG3A is a block diagram of a sonic training device according to an embodiment of the present invention;

[0022] FIG3B is a block diagram of an intelligent automation system according to an embodiment of the present invention;

[0023] FIG4A is a diagram of a received acoustic wave signal according to an embodiment of the present invention;

[0024] FIG4B is the physical change information (vertical direction) of the subject A according to an embodiment of the present invention; and

[0025] FIG. 4C shows the physical change information (horizontal direction) of the subject A according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] To make the above and / or other purposes, effects, and features of the present invention more clearly understood, preferred embodiments are described in detail below:

[0027] Please refer to Figure 1, which is a flow chart of a method according to one embodiment of the present invention. As shown in the figure, the acoustic wave annotation training method applied to physical change phenomena according to one embodiment of the present invention comprises the following steps:

[0028] Step S1: Capturing a physical change phenomenon with a capture unit;

[0029] Step S2: A transmitting end generates and transmits a corresponding acoustic wave signal according to the physical change phenomenon;

[0030] Step S3: receiving the sound wave signal with a receiving end and performing calculation analysis based on the change of the sound wave signal to generate corresponding physical change information; and

[0031] Step S4: A processing unit receives and labels the physical change phenomenon according to the physical change information to perform multimodal operation analysis and generate an analysis result.

[0032] As shown in step S1, the capture unit 1 captures a physical change phenomenon. In one embodiment, the type of the captured physical change phenomenon is selected from a detected electronic signal, image, audio and video, voice, text, or a combination of any two or more of the above. For example, the physical change phenomenon can be captured by capturing the image, or the physical change phenomenon can be directly presented by the electronic signal captured by the sensing unit, but the present invention is not limited thereto.

[0033] In one embodiment, the physical change phenomenon is selected from spatial position change, temperature change, elastic change, electromagnetic change, light change, wave change, vibration change, mass change, gas flow rate change, humidity change, or a combination of any two or more of the above, but is not limited thereto.

[0034] For example, the so-called physical change phenomenon can be a change in temperature, such as the process of boiling water or the process of burning, or it can be a change in the spatial position of a person / object / animal, such as the movement of a person / animal in space, or the movement of a robotic arm / robot or a vehicle, or it can be a spring pulling process, or it can be a change in light in the environment, or it can be a change in sound in the environment, or it can be a change in gas flow rate / humidity / temperature in the environment, or it can be a change in vibration of a machine or other object, or it can be a wave process of ebb and flow, and, among them, the physical change phenomenon can be one or more, for example: the image simultaneously contains the movement of a person / object / animal in space, temperature changes in space, light changes and sound changes, but is not limited to this.

[0035] As shown in step S2, the transmitter T1 can generate and transmit corresponding sound wave signals based on the physical change phenomena described above. In one embodiment, the transmitter T1 can be selected from a sound wave generating device, but is not limited to this. Any device with a speaker can be used, but is not limited to this. The frequency range of the sound wave signal can be between 20 Hz and 2 MHz.

[0036] In one embodiment, when there are multiple physical change phenomena, they can be transmitted through sound wave signals of different frequencies. For example, changes in human movement can be selected from 18kHz sound wave signals, changes in cat movement can be selected from 19kHz sound wave signals, changes in vehicle movement can be selected from 20kHz sound wave signals, changes in dog movement can be selected from 21kHz sound wave signals, changes in gas flow rate can be selected from 22kHz sound wave signals, and changes in temperature can be selected from 23kHz sound wave signals, but the present invention is not limited thereto.

[0037] As shown in step S3, the sound wave signal emitted in the previous step is received by the receiving end T2, and calculation analysis is performed based on the changes in the sound wave signal to generate corresponding physical change information. In one embodiment, the receiving end T2 can be selected from a smart phone, a tablet computer, a notebook computer, a personal computer, a television or a server, or can be from a headset, as long as it is a device with a microphone, but not limited to this. The corresponding physical change information can be further analyzed through the intensity changes of sound wave signals of different frequencies. For example, the distance value between the transmitting end T1 and the receiving end T2 can be analyzed by using the intensity changes of the sound wave signal to analyze the user's movement status, or the corresponding temperature value can be analyzed based on the intensity changes of the sound wave signal through an algorithm, or other similar analysis methods.

[0038] In one embodiment, the intensity change of the sound wave signal can be obtained based on a Fast Fourier Transform (FFT) operation. That is, the farther the distance, the lower the intensity. Conversely, the closer the distance, the higher the intensity. In this way, the change in the distance between the transmitter T1 and the receiver T2 can be inferred based on the intensity change. Furthermore, the frequency shift information can be calculated through the Doppler effect to generate a corresponding acceleration, thereby obtaining the direction of the velocity and the speed of the change in magnitude, but the present invention is not limited to this.

[0039] As shown in step S4, the processing unit 2 can annotate the physical change phenomenon based on the physical change information, thereby performing multimodal computing analysis to generate analysis results. For example, when the physical change phenomenon is a video of a person cooking a dish, the process of the person / spatula displacement and the process of the temperature change in the pot can be split into multiple images one by one, and each displacement / temperature change can be annotated in the corresponding image. In this way, the physical change phenomenon in the image can be automatically annotated in the corresponding image, achieving more accurate prediction, but not limited to this.

[0040] In one embodiment, the physical change information generated can be first converted by the processing unit 2 to generate corresponding semantic information. For example, when the physical change information is distance change and acceleration, the numerical changes in distance change and acceleration can be converted into text, for example, describing how the position of object A changes (moving from position A to position B), how the speed changes (how much distance object A moves per second increases to how much distance it moves per second), and how the direction changes (moving horizontally forward from position A to position B). Or when the physical change information is temperature change, the numerical change in temperature is converted into text, the temperature value is 100 degrees, or even when a specific temperature value is reached, its specific phenomenon can be described in text, such as: boiling state, but not limited to this.

[0041] In one embodiment, multimodal operation analysis includes but is not limited to data fusion analysis operations, multimodal learning and matching analysis operations. Multimodal operation analysis can combine different types of data, extract features from different types of data separately, and fuse the features. It can also match the features of different types of data with each other to improve the accuracy of the model. Preferably, ViT (Vision Transformer) can be used as a visual encoder to capture image features with physical change phenomena, and use a text encoder to match the physical change information with the image features with physical change phenomena to generate analysis results, but is not limited to this.

[0042] In one embodiment, after the analysis results are established, the corresponding prediction results can be generated through the analysis results through images / audio / speech / text. For example, the change in temperature value, distance value, or other physical phenomenon can be predicted through audio and video, or the corresponding prediction result can be the corresponding image / audio based on the description of the physical change phenomenon in the text / speech, but is not limited to this.

[0043] To demonstrate the beneficial effects of the present invention, training and prediction were performed using different physical phenomena, as described below:

[0044] Taking the water heating scenario as an example of the change in physical phenomena, Model 1 uses the acoustic wave training method of the present invention as image annotation and is trained with a multimodal model; Model 2 uses the predicted temperature value provided by the visual model as image annotation and is trained with a multimodal model; Model 3 uses the actual measured temperature as image annotation and is trained with a multimodal model; Model 4 manually groups images according to temperature range and uses classification-based machine learning to train images in different groups as training sets for the visual model.

[0045] Among the four models, Model 4 has high accuracy, but the visual model it uses cannot adapt to different background environments, and requires a lot of manual intervention in the integration of training set information. Model 3 needs to integrate temperature information and image time for image annotation data, which obviously requires a lot of manual intervention and needs to be checked one by one to avoid misjudgment. Model 2's data uses the current visual model as pre-processing. Although it is faster, it obviously cannot provide effective accuracy. Model 1 has both image and audio files during recording, and can automatically annotate the temperature information of the image at each time point directly and immediately through the program code, requiring almost no data pre-processing and annotation time. The table below shows that this acoustic wave-assisted image annotation of temperature information has great potential.

[0046] The training results are shown in Table 1 below:

[0047] Taking a falling ball as an example of a physical phenomenon, the actual total distance the ball free-falls is 1.65 meters. The sound wave annotation training method of the present invention uses Doppler's law to calculate the acceleration of sound frequency drift to provide image annotation training. The sound wave frequency drift calculation uses the final sound wave frequency drift to calculate the acceleration and then derive the distance and maximum speed. The inertia measurement unit calculation uses the inertia measurement unit data from the experiment to calculate the actual distance and maximum speed.

[0048] The predicted free fall distance and speed are shown in Table 2:

[0049] All data is calculated using the final image of a free-falling object touching the ground, using acoustic frequency drift, inertial measurement units, and the acoustic annotation training method of this invention. Acoustic frequency drift provides the closest real-world distance and speed calculations, and the analysis results from acoustic annotation training enable the prediction of object movement distance and speed. Furthermore, acoustic frequency drift accurately calculates the gravitational acceleration of a real-world free-falling object under varying volume and air resistance conditions, and provides the most convenient data pre-processing and annotation, significantly improving training efficiency, model iteration, and algorithm optimization.

[0050] Please refer to Figure 2A, which is a system diagram of one embodiment of the present invention. As shown in the figure, the acoustic wave annotation training system for physical change phenomena according to one embodiment of the present invention includes: an acquisition unit 1, a conversion unit T, and a processing unit 2. The processing unit 2 is signal-connected to the acquisition unit 1 and the conversion unit T, respectively, and is described in detail as follows:

[0051] The capture unit 1 is used to obtain physical change phenomena, wherein the physical change phenomenon can be presented in the form of images / audio / voice / text, or by a series of detected electronic signals, but is not limited thereto. In one embodiment, the capture unit 1 includes but is not limited to an image capture device, a voice capture device, an input device, a sensing device, or a combination of any two of the above, and can be configured according to needs. When the capture unit 1 is selected from the image capture device, the capture unit 1 captures image information and obtains the physical change phenomenon from the image information.

[0052] The conversion unit T includes a transmitting end T1 and a receiving end T2, whereby the transmitting end T1 generates and transmits a corresponding sound wave signal according to the physical change phenomenon, wherein the sound wave signal can be between 20 Hz and 2 MHz, that is, the transmitted sound wave signal can be a human audible signal or an ultrasonic signal, and the receiving end T2 receives this sound wave signal, and performs calculation analysis on the changes in the sound wave signal received by the receiving end T2 to generate corresponding physical change information, but is not limited to this.

[0053] In one embodiment, changes in the sound wave signal can be obtained through changes in the intensity / frequency / time of the corresponding received sound wave signal. In other words, when a physical phenomenon changes, a sound wave signal of a specific intensity / frequency / time can be transmitted through the transmitting end T1 and received by the receiving end T2. In this way, the intensity / frequency / time of the received sound wave signal will vary corresponding to the change in the physical phenomenon, and corresponding physical change information can be generated.

[0054] The processing unit 2 labels the corresponding physical change phenomenon according to the received physical change information, thereby performing multimodal operation analysis to generate an analysis result. The analysis result can be used to predict the corresponding physical change phenomenon.

[0055] In one embodiment, the operation processing unit 2 may further include a semantic conversion unit to generate corresponding semantic information after performing operation processing on the physical change information, so that the operation processing unit 2 can label the physical change phenomenon according to the semantic information and perform multimodal operation analysis, but this is not limited to this.

[0056] In one embodiment, please also refer to Figure 2B, which illustrates a schematic diagram of the system operation according to one embodiment of the present invention. As shown in the figure, taking the motion of a robotic arm as an example, the capture unit 1 captures the robotic arm's motion as it moves an object O from the left side to the right side. Simultaneously, the transmitter T1 transmits a corresponding acoustic signal based on the robotic arm's operating state. The receiver T2 receives the corresponding acoustic signal in real time, generating corresponding physical change information. This physical change information may include, but is not limited to, the spatial position change and the sound generated by the robotic arm during operation.

[0057] At this time, the processing unit 2 can mark the corresponding physical change phenomenon according to the physical change information, thereby performing subsequent multimodal operation analysis to generate corresponding analysis results. This analysis result can include the overall operating status of the robotic arm, which can be used as a result for detecting abnormal movements and abnormal noises of the robotic arm in the future, or can be used as actual robotic arm motion monitoring, for example: determining whether the direction of operation, rotational movement, or linear movement is correct, or determining whether there are other abnormal conditions.

[0058] Please refer to Figure 3A, which is a block diagram of a sound wave annotation training device according to one embodiment of the present invention. As shown in the figure, the sound wave annotation training device for physical change phenomena according to one embodiment of the present invention includes: a detection device 3 and a computing device 4. Detection device 3 is signal-connected to computing device 4 and is described in detail as follows:

[0059] The detection device 3 is used to generate and transmit sound wave signals based on physical change phenomena. In one embodiment, the detection device 3 includes but is not limited to the sensing device 31, wherein the detection device may not be provided with the sensing device 31, that is, it can directly generate and transmit corresponding sound wave signals based on the physical change phenomenon, or it can use the sensing device 31 to detect the physical change phenomenon, generate corresponding physical change information, and generate and transmit corresponding sound wave signals based on this physical change information. In other words, the transmitted sound wave signal itself contains physical change information, wherein the frequency range of the sound wave signal can be between 20Hz and 2MHz, but is not limited thereto.

[0060] In one embodiment, the so-called physical change phenomenon can be selected from spatial position change, temperature change, elastic change, electromagnetic change, light change, wave change, vibration change, mass change, gas flow rate change, humidity change, or a combination of any two or more of the above, but is not limited to this.

[0061] The computing device 4 is used to receive the sound wave signal emitted by the detection device 3, and thereby perform computational analysis based on the changes in the sound wave signal to generate corresponding physical change information. Thereby, the corresponding physical change phenomenon is labeled according to the physical change information and multimodal computational analysis is performed to generate analysis results. In one embodiment, the computing device 4 may include an input device, which is used to receive the sound wave signal emitted by the detection device 3. The input device may be one or more microphones, but is not limited to this.

[0062] In one embodiment, this system can also be used as a training device in an intelligent automation system. For example, while conventional intelligent robots or other autonomous devices are equipped with maps and positioning systems, they still cannot quickly grasp changes in relative spatial position when identifying their environment. Please also refer to FIG3B , which is a block diagram of an intelligent automation system according to one embodiment of the present invention. As shown in the figure, the intelligent automation system according to one embodiment of the present invention includes: an automation device 5, a detection device 6, a computing device 7, and a decision-making device 8. The automation device 5 is signal-connected to the detection device 6, the computing device 7 is signal-connected to the detection device 6, and the decision-making device 8 is signal-connected to both the automation device 5 and the computing device 7. The details are as follows:

[0063] The automated equipment 5 is a machine or system used to automatically perform specific tasks. In one embodiment, the automated equipment is selected from a robotic arm, a self-propelled device, a drone, or a robot, but is not limited thereto. Any equipment that can automatically perform specific tasks based on detection / identification results will suffice.

[0064] The implementation of the detection device 6 and the computing device 7 is the same as that of the previous embodiment, and thus will not be described in detail here.

[0065] The decision-making device 8 is used to predict and output decision information based on the analysis results generated by the computing device 7, thereby enabling the automation device 5 to operate according to the decision information. In other words, the intelligent automation system of the present invention transmits its physical phenomenon changes (such as spatial position changes, temperature changes, etc.) in real time through acoustic signals, thereby quickly completing positioning and subsequent movement or analysis for the intelligent automation system. In this way, it can avoid the need for traditional intelligent robots or other autonomous devices to spend a lot of time updating all information about the new environment. At the same time, it can adapt to environmental changes in various venues more quickly, but not limited to this.

[0066] To more clearly illustrate the embodiments of the present invention, examples are given below:

[0067] In one embodiment, please refer to Figure 4A, which illustrates a received acoustic signal diagram according to one embodiment of the present invention. As shown, an image capture device captures images of different subjects as they change in spatial position. Subject A is equipped with an 18kHz acoustic wave generator, while subject B is equipped with a 20kHz acoustic wave generator. The spatial changes of subjects A and B are reflected by the acoustic wave generators, which emit corresponding acoustic signals. A computing device 4 then calculates the corresponding spatial position and motion state based on the intensity changes of the received acoustic signals. The figure illustrates acoustic signals of different frequencies received by different subjects.

[0068] Please refer to Figures 4B and 4C, which are physical change information (vertical direction / horizontal direction) of subject A according to an embodiment of the present invention. As shown in the figure, the left side of Figure 4B indicates that subject A moves in the vertical direction relative to the image capture device, while the right side is the corresponding amplitude intensity change curve. The X-axis is time, and the Y-axis is the amplitude value obtained after Fourier transformation (i.e., the received sound wave signal intensity). The higher the amplitude value, the closer the subject A is to the image capture device. On the contrary, the lower the amplitude value, the farther the subject A is from the image capture device. Similarly, the left side of Figure 4C indicates that subject A moves in the horizontal direction relative to the image capture device, while the right side is the corresponding amplitude intensity change curve. In the intensity change curve, the X-axis represents time, and the Y-axis represents the amplitude value obtained after Fourier transformation (i.e., the intensity of the received sound wave signal). Similarly, the higher the amplitude value, the closer the distance between subject A and the image capture device is. Conversely, the lower the amplitude value, the farther the distance between subject A and the image capture device is. In this way, it can be known that subject A is located on the left or right side of the image capture device in the horizontal direction. Among them, the implementation method of subject B is the same as that of subject A, and the only difference is the sound wave frequency, so it is not repeated here.

[0069] In summary, the present invention provides a sound wave annotation training method, system, and equipment for physical change phenomena. The system uses sound wave signals as image annotations, thereby simplifying and automatically annotating images, reducing the degree of manual intervention, and solving the problem that physical change phenomena are difficult to train in images / audio. At the same time, it improves the accuracy of multimodal computing analysis, thereby achieving the purpose of the present invention.

[0070] The above is only a preferred embodiment of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and modifications without departing from the concept of the present invention. These improvements and modifications should also be considered within the scope of the present invention.

Claims

1. A sound wave annotation training method applied to physical change phenomena, characterized in that: The steps include: Acquiring a physical change phenomenon with a capture unit; A transmitting end generates and transmits a corresponding acoustic wave signal according to the physical change phenomenon; A receiving end receives and performs calculation analysis based on the change of the sound wave signal to generate corresponding physical change information; and A processing unit receives and labels the physical change phenomenon according to the physical change information to perform multimodal operation analysis and generate an analysis result.

2. The acoustic wave annotation training method for physical change phenomena according to claim 1, characterized in that: In the step of an operation processing unit labeling the physical change phenomenon according to the physical change information to perform multimodal operation analysis and generate an analysis result, the operation processing unit receives and performs operation processing according to the physical change information to generate corresponding semantic information, so that the operation processing unit labels the physical change phenomenon according to the semantic information.

3. The acoustic wave annotation training method for physical change phenomena according to claim 1, characterized in that: The type of the physical change phenomenon obtained is selected from detected electronic signals, images, audio and video, voice, text, or a combination of any two or more of the above.

4. The acoustic wave annotation training method for physical change phenomena according to claim 1, characterized in that: The physical change phenomenon is selected from spatial position change, temperature change, elastic change, electromagnetic change, light change, wave change, vibration change, mass change, gas flow rate change, humidity change or a combination of any two or more of the above.

5. A sound wave annotation training system applied to physical change phenomena, characterized in that: Include: a capture unit for capturing a physical change phenomenon; a conversion unit comprising a transmitting end and a receiving end, wherein the transmitting end is used to generate and transmit a corresponding acoustic wave signal according to the physical change phenomenon, and the receiving end is used to receive and perform calculation analysis based on the change of the acoustic wave signal to generate corresponding physical change information; and A processing unit is signal-connected to the capture unit and the conversion unit respectively, and is used for receiving and marking the physical change phenomenon according to the physical change information and performing multi-modal operation analysis to generate an analysis result.

6. The acoustic wave annotation training system for physical change phenomena according to claim 5, characterized in that: The capture unit is selected from an image capture device, a voice capture device, an input device, a sensing device, or a combination of any two or more of the above.

7. A sound wave annotation training device applied to physical change phenomena, characterized in that: Include: A detection device for generating and emitting an acoustic wave signal in response to a physical change phenomenon; and A computing device is signal-connected to the detection device, and is used to receive and perform computing analysis based on the changes in the acoustic signal to generate corresponding physical change information, and annotate the physical change phenomenon based on the physical change information to perform multimodal computing analysis and generate an analysis result.

8. The acoustic wave annotation training device for physical change phenomena according to claim 7, characterized in that: The detection device includes a sensing device for detecting the physical change phenomenon and generating the physical change information.

9. An intelligent automation system, characterized in that: Include:

1. Automation equipment; a detection device, connected to the automation device by signal, for acquiring image information and generating and emitting a corresponding acoustic wave signal according to a physical change phenomenon in the image information; a computing device, signal-connected to the detection device, configured to receive and perform computational analysis based on changes in the acoustic signal to generate corresponding physical change information, and annotate the image information based on the physical change information to perform multimodal computational analysis and generate an analysis result; and A decision-making device is respectively connected to the automation device and the computing device via signals, and is used to predict and output decision information according to the analysis result, so that the automation device operates according to the decision information.

10. The intelligent automation system according to claim 9, characterized in that: The automated equipment is selected from a robotic arm, a self-propelled device, a drone or a robot.

Citation Information

Patent Citations

  • Acoustic detection method and device

    CN109212482A

  • Positioning method, device and system of intelligent equipment, intelligent equipment and storage medium

    CN113075618A

  • Gas concentration monitoring device based on sound wave generator

    CN116660366A

  • Ultrasonic Test Equipment and Evaluation Method Thereof

    US20140230556A1