INTRODUCTION
Among the
biggest challenges in teaching a foreign language is ensuring that students are
fully prepared to use the target language. In communication, one skill that has
a vital role is speaking. However, this is not easy to do, considering the many
aspects that need to be learned in speaking skills, including pronunciation.
Speaking and pronunciation have an inseparable relationship. Communication can
be disrupted or even misunderstood if the pronunciation of the message is not
articulated
Aligned
with the importance of pronunciation, the main problem discussed in this study
is that many students experience persistent pronunciation errors, which hinder
practical and clear communication. Students who make persistent pronunciation
mistakes make them feel insecure about their communication skills. Although the
teacher has provided pronunciation examples regularly, many students still
struggle to pronounce words correctly. This difficulty is partly due to the
limited class time and the lack of practice given. This indicates the need for
practice aids that use technology to facilitate independent practice for
students. In addition, students often do not receive immediate or personalized
feedback on their mistakes in traditional learning that does not utilize
technology, making it difficult to correct this problem effectively. In
addition, some other factors still cause this problem, such as their mother
tongue, their learning environment, and lack of knowledge about English
phonology and phonetics
Speechling,
as an AI-supported tool, offers a promising alternative. Previous research by
García-Sánchez et al. (2020) and Zou et al. (2021) shows that AI-based
pronunciation feedback improves pronunciation skills among non-native English
speakers. Additionally, studies like Li and Li (2021) demonstrate how AI
technology, through personalized practice and instant feedback, can enhance
learners' confidence. Speechling is also one of the technologies based on
Automatic Speech Recognition (ASR). ASR technology can provide corrective
feedback on pronunciation in different way
The
theoretical framework is based on Computer-Assisted Language Learning (CALL)
theory, which emphasizes the role of technology in enhancing language learning
and highlights the utilization of the Speechling website. There are suggestions
for students to practice longer with ASR-based language learning systems to
achieve high pronunciation performance
Therefore, AI-supported CALL techniques can customize the language learning process to accommodate varying skill levels among students, optimizing it in a more flexible and adaptable manner. This study investigates whether providing easily accessible mutual practice and quick feedback within the Speechling website can improve students' pronunciation. This research has an important role as it will address the need for a practical digital tool to help improve students' pronunciation. If Speechling is shown to effectively assist students in correcting their pronunciation, it will be a handy tool for educators and curriculum developers and can be implemented into lessons. In addition, it can also provide opportunities for students to learn independently, improve their pronunciation, and boost their confidence.
LITERATURE REVIEW
Based
on the review of previous studies, this study will focus on two areas.
1. Automatic
Speech Recognition (ASR) for Pronunciation Learning.
Automatic
speech recognition (ASR) is a machine-based, autonomous method for decoding and
recording spoken language (Levis & Suvorov, 2012). Distinguishing
between speech recognition and speech understanding (or speech identification)
is essential, where speech understanding is the process of interpreting the
meaning of an expression rather than simply transcribing the words. Speech
recognition is different from voice recognition. Speech recognition is related
to a machine's ability to identify spoken words (i.e., what is being said). In
contrast, speech recognition is associated with a machine's ability to
recognize a person's speaking style (i.e., who is speaking the words).
According to (Levis & Suvorov, 2012), Speech recognition systems can
be distinguished based on three main things: speaker dependency, speech
continuity, and vocabulary size. Based on the speech data used for training,
ASR systems can be speaker-dependent (where the system needs to be trained for
each speaker), speaker-independent (where the system can recognize new speakers
because its training data comes from many speakers), or adaptive (where the
system starts without recognizing a particular speaker, but then adapts to a
specific speaker through training). Based on speech continuity, there are (a)
isolated word recognition systems, which recognize words spoken fragmentarily;
(b) connected word recognition systems, which can recognize words spoken
without pauses; (c) continuous speech recognition systems, which can recognize
complete sentences without pauses between words; and (d) word scanning systems,
which retrieve words or phrases from continuous speech. Finally, ASR systems
can also be differentiated based on the size of the vocabulary used in
training, whether small or large.
Automatic
Speech Recognition (ASR) technology has been suggested in recent studies on
technological developments as a valuable tool to enhance learners' FL speaking
skills. Cucchiarini et al. (2009) found that the feedback provided
was compelling enough to help students enhance their pronunciation after
practice in a relatively short period, even though the developed ASR system has
not achieved perfect accuracy in detecting students' pronunciation errors.
Research by McAndrews (2020) supports the use of technology in language
learning, as it has proven effective in improving students' pronunciation
skills, primarily in prosody. Researchers have started using ASR technologies
that convert speech into text freely, such as Google Speech Recognition (GSR)
(e.g., Mroz, 2018) for spoken exercises to improve pronunciation skills.
Previous research has generally found that learners have a favorable view of
pronunciation training using ASR, as this technology allows them to practice
independently (McCrocklin, 2016) and provides direct feedback
(Inceoglu, 2022; Liakin et al., 2017).
2.
Teaching Pronunciation
Many
aspects must be considered when speaking, one of them being pronunciation.
Pronunciation has an essential role in communication, as when someone does not
pronounce words clearly, it will undoubtedly cause the meaning to be conveyed,
not delivered. Good pronunciation is necessary for effective communication,
although other language skills are also essential (Alghazo, 2015). Teaching
pronunciation holds significant importance as it directly impacts the
interpretation of sentences, typically conveyed through both segmental and
suprasegmental elements. Segmental features pertain to the smallest phonetic
sound units (Pennington & Richards, 1986, p. 208). Proper pronunciation
teaching is essential to reduce misunderstandings in communication, especially
in foreign languages. Thus, pronunciation teaching should focus on helping
students develop clear and understandable speech (Alghazo et al., 2023).
For several reasons, teaching proper pronunciation is essential in
communication, especially in the scope of foreign languages. First,
pronunciation plays a significant role in achieving effective communication.
Adequate pronunciation can help ensure that the recipient of the message can
understand the sentence we convey because if there are errors in pronunciation,
it can hinder the communication process because of mistakes in interpreting
meaning. Secondly, proper pronunciation also affects someone's listening
ability. Learners who pronounce and recognize sounds correctly find it easier
to understand the original speaker. Third, adequate pronunciation indicates
that someone has good language skills and can help increase learners'
confidence. (Alsuhaibani et al., 2024). Pronunciation teaching has a very
significant role in improving the accuracy of language production among L2/FL
learners of various languages. (Piske et al., 2001).
CLASSROOM APPLICATION
After
introducing the Speechling website, teachers can guide students step by step on
how to use it in class. Then, students are given time to listen to recordings
of native speakers from Speechling. After that, students can be asked to try to
imitate and record their voices. Teachers monitor students' learning progress
using the Speechling website and help them with difficult words or
pronunciations.
Students
are encouraged to continue using Speechling outside of class as a daily
exercise. Teachers can also give students simple assignments, such as recording
a few sentences daily to help them develop consistent habits. Due to limited
time in class, this can help students get more practice.
Teachers
can conduct pronunciation tests both before and after using Speechling to
monitor students' pronunciation progress, which can help measure improvements
in their pronunciation skills. Additionally, teachers can provide reflection
sheets where students can share their experiences using the tool, what they
have learned, and any challenges they have faced.
CHALLENGES AND OPPORTUNITIES
Challenges
1. Limited Access to Technology
Not all students have smartphones, laptops, or stable internet connections at home, so some students will struggle to practice using Speechling outside of class.
2. Low Student Motivation
Some students may not be motivated to practice pronunciation regularly. When they are tired, bored, or lack confidence in their speaking abilities, this will further demotivate them from practicing pronunciation regularly.
3. Basic Technology Skills
Students who are less tech-savvy or
unfamiliar with websites or apps may initially have difficulty using
Speechling. They may need additional guidance before they can use this tool
effectively.
Opportunities
1. Flexible Practice Time
Students can use Speechling anytime and anywhere, giving them more opportunities to improve their pronunciation since class time is limited.
2. Instant Feedback for Learning
Speechling provides quick feedback on how they are pronouncing words. This helps students identify their mistakes quickly and try to correct their pronunciation until it is correct.
3. Increased Confidence and Independence
By seeing their progress through
regular practice, students will become more confident and take greater
responsibility for their learning.
REFERENCES
Alghazo, S., Jarrah, M., & Al Salem, M. N. (2023). The
efficacy of the type of instruction on second language pronunciation
acquisition. Frontiers in Education, 8. https://doi.org/10.3389/feduc.2023.1182285
Alsuhaibani, Y., Mahdi, H. S., Al khateeb, A., Al Fadda, H. A.,
& Alkadi, H. (2024). Web-based pronunciation training and learning
consonant clusters among EFL learners. Acta Psychologica, 249.
https://doi.org/10.1016/j.actpsy.2024.104459
Bashori, M., Hout, R. Van, Strik, H., & Cucchiarini, C.
(2024). I Can Speak : improving English pronunciation through automatic speech
recognition-based language learning systems I Can Speak : improving English
pronunciation through automatic. Innovation in Language Learning and
Teaching, 1–19. https://doi.org/10.1080/17501229.2024.2315101
Cucchiarini, C., Neri, A., & Strik, H. (2009). Oral
proficiency training in Dutch L2: The contribution of ASR-based corrective
feedback. Speech Communication, 51(10), 853–863. https://doi.org/10.1016/j.specom.2009.03.003
Inceoglu, S., Chen, W.-H., &
Lim, H. (2024). Monitoring student behavior in autonomous automatic speech
recognition-based pronunciation practice. System, 124, 103387. https://doi.org/10.1016/j.system.2024.103387
Khalid, F., & Almuslimi, A. (2020). Pronunciation Errors
Committed by EFL Learners in the English Department in Faculty of Education -
Sana ’ a University.
Levis, J., & Suvorov, R. (2012). Automatic Speech
Recognition. In The Encyclopedia of Applied Linguistics. Wiley.
https://doi.org/10.1002/9781405198431.wbeal0066
Loi, D. P. (2018). An Analysis of Vietnamese EFL Students’
Pronunciation of English Affricates and Nasals. International Journal of
English Linguistics, 8(2), 298.
https://doi.org/10.5539/ijel.v8n2p298
McCrocklin, S. M. (2016). Pronunciation learner autonomy: The
potential of Automatic Speech Recognition. System, 57, 25–42.
https://doi.org/10.1016/j.system.2015.12.013
Novianti, N. (2021). Introducing Critical Literacy to Pre-Service
English Teachers through Fairy Tales. Journal of Literary Education, 4,
196–215. https://doi.org/10.7203/jle.4.21026