Building a National Corpus

Building a National Corpus
Title Building a National Corpus PDF eBook
Author Dawn Knight
Publisher Springer Nature
Pages 192
Release 2021-10-08
Genre Language Arts & Disciplines
ISBN 3030818586

Download Building a National Corpus Book in PDF, Epub and Kindle

This book aims to provide a micro-level, working model of a methodological approach and practical guidelines for building a corpus, informed by the work on the CorCenCC project (Corpws Cenedlaethol Cymraeg Cyfoes - the National Corpus of Contemporary Welsh). It focuses specifically on the development of detailed design frames for corpora across communicative modes (spoken, written and e-language), and the practical processes involved in the planning, collection, transcription, collation and (re)presentation of language data. The book is designed to be of significant value and relevance to those interested in critically engaging with corpus methodology. Although Welsh is the language under discussion, the processes and approaches discussed in the building of CorCenCC can be applied to a lesser or greater extent to other language contexts. This book provides a working model, and an account of how to build a corpus dataset from which step by step guidelines for creating other linguistic corpora in any language can be easily extrapolated. It will be of value to students and scholars of minority languages and corpus linguistics.

Developing Linguistic Corpora

Developing Linguistic Corpora
Title Developing Linguistic Corpora PDF eBook
Author Martin Wynne
Publisher Oxbow Books Limited
Pages 100
Release 2005
Genre Language Arts & Disciplines
ISBN

Download Developing Linguistic Corpora Book in PDF, Epub and Kindle

A linguistic corpus is a collection of texts which have been selected and brought together so that language can be studied on the computer. Today, corpus linguistics offers some of the most powerful new procedures for the analysis of language, and the impact of this dynamic and expanding sub-discipline is making itself felt in many areas of language study. In this volume, a selection of leading experts in various key areas of corpus construction offer advice in a readable and largely non-technical style to help the reader to ensure that their corpus is well designed and fit for the intended purpose. This guide is aimed at those who are at some stage of building a linguistic corpus. Little or no knowledge of corpus linguistics or computational procedures is assumed, although it is hoped that more advanced users will find the guidelines here useful. It is also aimed at those who are not building a corpus, but who need to know something about the issues involved in the design of corpora in order to choose between available resources and to help draw conclusions from their studies.

Overcoming Challenges in Corpus Construction

Overcoming Challenges in Corpus Construction
Title Overcoming Challenges in Corpus Construction PDF eBook
Author Robbie Love
Publisher Routledge
Pages 183
Release 2020-01-06
Genre Language Arts & Disciplines
ISBN 0429771096

Download Overcoming Challenges in Corpus Construction Book in PDF, Epub and Kindle

This volume offers a critical examination of the construction of the Spoken British National Corpus 2014 (Spoken BNC2014) and points the way forward toward a more informed understanding of corpus linguistic methodology more broadly. The book begins by situating the creation of this second corpus, a compilation of new, publicly-accessible Spoken British English from the 2010s, within the context of the first, created in 1994, talking through the need to balance backward capability and optimal practice for today’s users. Chapters subsequently use the Spoken BNC2014 as a focal point around which to discuss the various considerations taken into account in corpus construction, including design, data collection, transcription, and annotation. The volume concludes by reflecting on the successes and limitations of the project, as well as the broader utility of the corpus in linguistic research, both in current examples and future possibilities. This exciting new contribution to the literature on linguistic methodology is a valuable resource for students and researchers in corpus linguistics, applied linguistics, and English language teaching.

English Corpus Linguistics

English Corpus Linguistics
Title English Corpus Linguistics PDF eBook
Author Charles F. Meyer
Publisher
Pages 168
Release 2002
Genre Computational linguistics
ISBN 9780511044755

Download English Corpus Linguistics Book in PDF, Epub and Kindle

English Corpus Linguistics is a step-by-step guide to creating and analyzing linguistic corpora. The author shows how to collect and computerize data for inclusion in a corpus; how to annotate the data; and how to conduct a linguistic analysis of it once it has been created.

Using Corpora in Discourse Analysis

Using Corpora in Discourse Analysis
Title Using Corpora in Discourse Analysis PDF eBook
Author Paul Baker
Publisher Bloomsbury Publishing
Pages 281
Release 2023-08-24
Genre Language Arts & Disciplines
ISBN 1350083771

Download Using Corpora in Discourse Analysis Book in PDF, Epub and Kindle

How can you carry out discourse analysis using corpus linguistics? What research questions should I ask? Which methods should you use and when? What is a collocational network or a key cluster? Introducing the major techniques, methods and tools for corpus-assisted analysis of discourse, this book answers these questions and more, showing readers how to best use corpora in their analyses of discourse. Using carefully tailored case studies, each chapter is devoted to a central technique, including frequency, concordancing and keywords, going step by step through the process of applying different analytical procedures. Introducing a wide range of different corpora, from holiday brochures to political debates, the book considers the key debates and latest advances in the field. Fully revised and updated, this new edition includes: - A new chapter on how to conduct research projects in corpus-based discourse analysis - Completely rewritten chapters on collocation and advanced techniques, using a corpus of jihadist propaganda texts and covering topics such as social media and visual analysis - Coverage of major tools, including CQPweb, AntConc, Sketch Engine and #LancsBox - Discussion of newer techniques including the derivation of lockwords and the comparison of multiple data sets for diachronic analysis With exercises, discussion questions and suggested further readings in each chapter, this book is an excellent guide to using corpus linguistics techniques to carry out discourse analysis.

Statistics in Corpus Linguistics

Statistics in Corpus Linguistics
Title Statistics in Corpus Linguistics PDF eBook
Author Vaclav Brezina
Publisher Cambridge University Press
Pages 317
Release 2018-09-20
Genre Foreign Language Study
ISBN 1107125707

Download Statistics in Corpus Linguistics Book in PDF, Epub and Kindle

A comprehensive and accessible introduction to statistics in corpus linguistics, covering multiple techniques of quantitative language analysis and data visualisation.

Corpus Linguistics and Linguistically Annotated Corpora

Corpus Linguistics and Linguistically Annotated Corpora
Title Corpus Linguistics and Linguistically Annotated Corpora PDF eBook
Author Sandra Kuebler
Publisher Bloomsbury Publishing
Pages 321
Release 2014-12-18
Genre Language Arts & Disciplines
ISBN 1441119809

Download Corpus Linguistics and Linguistically Annotated Corpora Book in PDF, Epub and Kindle

Linguistically annotated corpora are becoming a central part of the corpus linguistics field. One of their main strengths is the level of searchability they offer, but with the annotation come problems of the initial complexity of queries and query tools. This book gives a full, pedagogic account of this burgeoning field. Beginning with an overview of corpus linguistics, its prerequisites and goals, the book then introduces linguistically annotated corpora. It explores the different levels of linguistic annotation, including morphological, parts of speech, syntactic, semantic and discourse-level, as well as advantages and challenges for such annotations. It covers the main annotated corpora for English, the Penn Treebank, the International Corpus of English, and OntoNotes, as well as a wide range of corpora for other languages. In its third part, search strategies required for different types of data are explored. All chapters are accompanied by exercises and by sections on further reading.